Vizard Agent

How to get a text-free cut you will caption yourself

Last updated 2026-09-22 · 8 min read

Tell Vizard Agent that the cut must contain zero text and that you are adding captions yourself afterwards. Both halves matter: the first rules out titles, labels and logos, and the second tells it to carry emphasis through framing and cutaways instead, and to leave the space your captions will need.

What is the short version?

You want a properly edited cut, but the captions are yours — your fonts, your timing, your style, done in your own tool. What you need back is a finished edit with nothing written on it. Vizard Agent treats that as a design constraint, not just an omission.

  1. Tell Vizard Agent zero text, and say where you will add yours.
  2. Expect emphasis to come from framing rather than labels.
  3. Check the crop leaves room for your captions.

What do you need before you start?

The footage, and a sense of how you caption. If your captions sit in the lower third, the cut should not put the speaker's hands there; if you use large karaoke text across the middle, the framing has to survive it. Telling Vizard Agent where your text will go changes the crop it chooses.

You should also decide about music. A text-free cut usually still wants a bed, and a bed mixed well under the voice is one less thing for you to do afterwards. Say if you want the audio finished or left flat.

What do you type into Vizard Agent?

List what counts as text. This sounds excessive and is not, because "no text" has edge cases that people disagree about — a logo, an emoji, a sticker, a lower third and a burnt-in subtitle are all text to you and not all of them read that way by default.

Edit this from beginning to end, but the final video must contain zero text: no subtitles, captions, titles, headings, overlays, emoji, stickers or logos. I will add all text myself in CapCut afterwards.

That was close to the real brief behind this article, list and all. It is worth the extra line, because finding a logo in the corner after you have captioned the whole thing is an expensive discovery.

What does Vizard Agent actually do?

It edits to the meaning of the speech, then carries the emphasis with picture rather than with words. Without text, the tools available are framing, timing and cutaways — a gentle push in on the sentence that matters, a cutaway to the thing being described, a slightly longer beat after a key point.

In that session the work ran like this:

That third step is the one people skip. On an explanatory video the cutaway has to be correct, not just relevant, and generated footage of a specialist subject is very often wrong in ways the audience will notice and you will not.

What does the result look like?

A finished-feeling edit with a completely clean frame. The speech flows, the filler is gone, the music sits under the voice, and there is not a pixel of type anywhere — which means it drops into your own project as a base layer rather than as something you have to work around.

Vizard Agent reviews the rendered frames specifically for stray text before delivery. A logo that arrived with a stock clip or a watermark on a cutaway is the failure mode here, and it is only visible if someone is looking for it.

When does this not work well?

When the explanation genuinely needs a label. Some points cannot be made by pointing the camera — a number, a name, a comparison between two things — and a text-free cut will leave those moments doing less work than they could. That is a real cost of the constraint, and it is yours to accept.

It also limits what can be done about a weak stretch. With text you can hold a viewer through a slow passage with an on-screen question; without it, the only options are cutting it or covering it with a cutaway. Vizard Agent will tell you where that trade came up.

And if your own captions end up covering the emphasis framing, the two treatments fight. Telling Vizard Agent where your text will sit avoids that, but only if you know before the cut is built.

How do you fix a result that came back wrong?

If there is text in it anywhere, say where you saw it. Vizard Agent finds the source — usually a stock clip, a cutaway or a watermark that came in with an asset — and replaces that element rather than masking it over.

If the emphasis is in the wrong place, name the sentence. Because the emphasis is carried by framing, moving it is a re-render of that section and not a rebuild, and Vizard Agent keeps the audio untouched while it does it.

If the crop does not suit your captions, say what your caption block looks like. Vizard Agent can re-frame to leave the space rather than you having to shrink your own text to fit.

How does Vizard Agent compare to doing it yourself?

You could do the cut yourself and have nothing to strip out afterwards, which is the honest comparison. The reason to split the job is that the cut and the captions are different skills, and most people doing this are confident in their captioning and less so in their pacing.

What Vizard Agent brings is the part before the text: finding the retakes, cutting the filler without cutting the meaning, and choosing the cutaway. Delivering that clean, and checking it is clean, is what makes the split work — a cut that arrives with someone else's design decisions baked in is worse than no cut at all.

Common questions

Does zero text include my logo? Yes, unless you say otherwise. Vizard Agent leaves it out and you add it with your captions.

Can I get the transcript too? Yes, and you should. Vizard Agent can deliver the word timings so your captions do not have to be typed.

Will there be music? Only if you ask. Vizard Agent can deliver the cut with a bed or with dialogue alone.

What about burnt-in text in my source footage? That is already there. Tell Vizard Agent whether to crop it out or leave it.

Can I ask for a version with text as well? Yes. Vizard Agent renders both from the same build rather than adding text afterwards.

How do I say where my captions go? Describe the block. Vizard Agent frames to leave that area clear.

Will the emphasis zooms fight my text? Not if you say where the text sits. Vizard Agent keeps the movement away from it.

Does the cutaway footage have a watermark? It should not. Vizard Agent checks delivered frames for stray marks.

Can I get it in CapCut's preferred format? Yes. Name the resolution and codec and Vizard Agent exports to it.

What if I want subtitles later from the same cut? Ask Vizard Agent for the timings now; they stay valid for the delivered cut.

Does removing filler change the length? Yes, and Vizard Agent reports the final duration so your caption plan matches.

Can it keep the original audio untouched? Yes. Say so and Vizard Agent copies the audio rather than processing it.

Will it add transitions? Only picture transitions, never text ones. Vizard Agent treats animated titles as text.

What if a point really needs a label? Vizard Agent flags the moment rather than adding one, so you can decide.