Vizard Agent

How to turn a video's captions into a reusable green-screen layer

Last updated 2026-09-17 · 8 min read

Ask Vizard Agent for a green-screen version containing only the text already on the video, and say whether the original letterforms matter. There are two ways to do it — isolate the real pixels, or retype the words in a new font — and they give very different results from the same request.

What is the short version?

You want the captions from a video as a layer you can drop onto something else. That means everything except the letters becomes a flat green, and the letters have to survive the process looking exactly as they did.

  1. Ask Vizard Agent for a green-screen version with only the text kept.
  2. Say whether the original font must be preserved.
  3. Say the audio stays, if you want to keep it.

What do you need before you start?

The video and a decision about fidelity. Vizard Agent can rebuild the words as clean new captions, which key perfectly and look slightly different, or isolate the original strokes, which keeps the typeface and brings a little background noise with it.

What do you type into Vizard Agent?

Say which of the two versions you want, or ask to see both. Left unspecified, a request for "the text on green" is ambiguous, and the version that keys most cleanly is also the version that has quietly replaced the design you were trying to reuse.

Prompt

Turn this video into a green screen that keeps only the text already on it — same font, same position, same timing. Keep the original audio. Prioritise keeping the real letter pixels over a perfectly clean key, and leave the watermark at the bottom out.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on a captioned clip that had to become a reusable text layer. It ends up delivering two different versions of the same thing, because the first one answers the request too tidily — clean, keyable, and set in a typeface nobody asked for.

  1. Extracts frames to see the real layout and identify each text layer.
  2. Builds a short green-screen sample to test whether the text separates at all.
  3. Corrects the colour filter expression and rebuilds the sample with a channel mask.
  4. Reviews the sample for green purity and leftover background.
  5. Surveys every text cue in the whole video so background detail is not kept by mistake.
  6. Narrows the mask to the central text region and the watermark strip.
  7. Reads the text and its timings with an analysis pass, as a route to rebuilt captions.
  8. Exports a clean green screen with regenerated text layers and the original audio.
  9. Adjusts the size and region of the white text to match the original layout.
  10. Separates the coloured text layers from the white line so they stop overlapping.
  11. Switches approach: isolates the original strokes directly onto the green.
  12. Tests a narrower mask to keep the original typeface with less background.
  13. Widens it again when the narrow mask eats parts of the letters.
  14. Delivers the version that keeps the most original text pixels, audio intact.

Steps eleven to fourteen are the interesting part. A rebuilt caption keys beautifully and is not the same object; once that became clear, Vizard Agent stopped improving the rebuild and went back to isolating what was actually filmed.

Steps twelve and thirteen are the trade-off in miniature. Tighten the mask and the background disappears along with the thin parts of the letters; loosen it and the strokes come back with a halo. Vizard Agent tests both directions rather than settling on the first acceptable answer.

Step five is what stops a bright object in the background being mistaken for a letter, which is the usual reason these layers have mysterious specks in them.

What does the result look like?

A video the same length as the original, flat green everywhere except the words, with the original typeface, position and timing. The audio is the source audio, and you can key it over anything — Vizard Agent checks the green for purity and the letters for missing strokes before delivering.

When does this not work well?

Some text cannot be lifted off its background. Letters over a busy scene, semi-transparent captions, or text with a soft shadow share their edge pixels with the picture behind them, and no key separates those cleanly — the honest route there is rebuilding the captions.

How do you fix a result that came back wrong?

Say whether you are seeing too much or too little of the original picture. Vizard Agent moves the mask one way or the other from there, and while you are deciding it is worth asking for a short sample rather than a full render each time.

How does Vizard Agent compare to doing it yourself?

By hand you would either key it in an editor, fighting the same trade-off with a luma threshold, or retype every caption and accept that the typeface is now approximate. Most people retype, because the key never quite comes clean and retyping at least looks deliberate.

By hand Vizard Agent
The key One threshold, adjusted by eye Tested tight and loose, compared
The font Usually replaced Preserved unless you say otherwise
Stray detail Keyed in with the text Text cues surveyed across the video
Delivery One attempt The rebuild and the isolation, both
Audio Often lost Kept and verified

Common questions

Why green rather than transparent? Either works. Ask Vizard Agent for an alpha channel if your tool prefers it.

Can it keep only some of the text? Yes. Name the lines or the region, and Vizard Agent excludes the rest.

Will the timing match the original? Yes. Vizard Agent keeps each line appearing and disappearing exactly where it did.

Can it keep the watermark? Yes or no, as you prefer — say which in the instruction.

Is a rebuild ever better? Yes, when the original text is low resolution or badly compressed.

Does it keep the audio? By default, and Vizard Agent verifies it on the delivered file.

Can I get a transparent WebM? Yes. Ask Vizard Agent for one and green spill stops being a concern.

What if the text moves? Vizard Agent isolates it frame by frame, so the layer follows the movement.

Will it work on handwriting? Sometimes. Thin strokes are the first thing a tighter mask loses.

How do I check the key? Put it over a dark background — spill and haloes show there first.

Can it also give me the text as a file? Yes, ask Vizard Agent for the transcript of the on-screen text as well.

Does it re-encode the whole video? Vizard Agent renders a new file and leaves your source untouched.

What about vertical videos? No difference. The frame shape does not change the approach.

Can it do several clips the same way? Yes, once the mask settings are agreed for that style of caption.