Vizard Agent

How to turn the captions off for the shots where the picture is the point

Last updated 2026-09-15 · 8 min read

Tell Vizard Agent which shots are the ones the viewer came to look at, and it removes the captions from those while keeping them everywhere else. It does this as a captions-only pass that leaves the edit and the audio untouched, then inspects each close-up at full resolution to confirm no text sits over it.

What is the short version?

Captions are on by default and they are right most of the time. They are wrong on the shot where you hold something up to the lens — because that shot is not illustrating what you are saying, it is the thing being said.

  1. Tell Vizard Agent which shots are the content itself.
  2. Ask for a captions-only change, nothing else.
  3. Ask for those shots checked at full size.

What do you need before you start?

The finished video and a way to describe the shots. Vizard Agent can find them by description — "whenever I hold a card up close" — which is usually easier than listing timecodes, and stays right if the edit changes later.

What do you type into Vizard Agent?

Describe the shots by what happens in them, and say that nothing else may change. The second half matters because removing captions from part of a video touches the caption track, and a caption track rebuilt from scratch tends to re-time everything.

Prompt

Keep this edit exactly as it is and change only the captions. Remove them whenever I am holding a card up close to the camera — those shots are what people are watching — and keep the captions during the intro and the bits between. Check the close-ups at full size so I can see nothing is sitting over them.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on a real unboxing video whose captions were sitting across the items being unboxed. The pass is deliberately narrow: the edit, the audio and the music bed are not rebuilt at any point, and only the caption layer is touched.

  1. Checks caption readability over the important shots at full resolution.
  2. Fixes captions running outside the vertical frame.
  3. Builds a captions-only version without touching the edit or the audio.
  4. Confirms captions now appear only in the intro and the linking passages.
  5. Inspects the close-ups up close to confirm they are free of text.
  6. Checks the on-screen labels are gone as well as the captions.
  7. Re-checks the structure and audio are unchanged by the pass.
  8. Prepares any new close-ups from the source with no text over them.
  9. Assembles the result on the existing base rather than rebuilding it.
  10. Reviews every close-up again in the finished file.

Step three is the whole discipline. A captions-only version means the thing you approved stays approved, and the only difference between the two files is the text layer.

Step five is not the same as watching it back. A caption at playback size can look like it is beside the object and, at full resolution, be sitting across the corner of it — which is exactly where the detail people are pausing for tends to be.

Step six catches what people forget. Captions are not the only text: a label you added earlier, a price tag, a name strip, all sit over the same shot and all need the same decision.

What does the result look like?

The same video with the text where it helps and absent where it does not. Vizard Agent gives you the close-ups at full resolution as the evidence, and confirms the edit, the audio and the running time are identical to the version you approved.

When does this not work well?

Captions that come and go can feel broken rather than deliberate. If the shots are short and frequent, the text will flicker on and off, and it is often better to move the captions than to remove them.

How do you fix a result that came back wrong?

Name the moment and say what is still sitting over it. Because Vizard Agent worked on the caption layer alone, every correction you ask for applies to that layer too, and the edit underneath cannot be disturbed by them however many rounds it takes.

How does Vizard Agent compare to doing it yourself?

By hand you delete the caption clips over the close-ups, and then the caption track is a different length, so the ones after it shift. You fix that, re-export, and discover at full resolution that two of them still clip the corner of the object.

By hand Vizard Agent
Finding the shots Scrubbing for them Found from a description
The rest of the edit At risk every time Untouched by a captions-only pass
Checking At playback size At full resolution
Other on-screen text Forgotten Cleared with the captions
Later close-ups Have to be redone Prepared without text from the start

Common questions

Can it find the shots itself? Yes. Describe what happens in them and Vizard Agent locates every instance.

Will the edit change? No. Vizard Agent builds a captions-only version from the same base.

Will the length change? No, and Vizard Agent confirms the duration against the previous export.

Can it move captions instead of removing them? Yes, and for frequent short shots Vizard Agent will suggest that instead.

What about text I added earlier? Tell Vizard Agent and it clears those from the same shots.

Does the audio get re-mixed? No. A captions-only pass leaves the mix exactly as approved.

How do I know nothing is over the object? Vizard Agent inspects each close-up at full resolution and shows you.

Can I keep captions for the speech over a close-up? Yes. Say which matters more and Vizard Agent follows that rule.

What if I add more close-ups later? Vizard Agent prepares new inserts with no text over them from the start.

Does this work in a vertical frame? Yes, and Vizard Agent also checks no caption runs outside the frame.

Can I get both versions? Yes. Ask Vizard Agent for the captioned and the cleaned version.

What about accessibility? Say if captions must be continuous; Vizard Agent will move rather than remove.

Will the caption style change? No. Only where they appear changes, not how they look.

Can it do this on several videos? Yes, with the same rule applied to each.