Vizard Agent

How to separate the two colours of subtitle burnt into a video

Last updated 2026-09-22 · 8 min read

Tell Vizard Agent that the video carries two kinds of burnt-in text in two colours and that you need them kept apart. It reads each colour in its own region of the frame, works out which one is dialogue and which is annotation, and returns a timed translation with the two tracks separated.

What is the short version?

Short-form video from some markets carries two text layers at once: the spoken line in white along the bottom, and a green note explaining what is happening. Transcribe them together and you get a document where a narrator's aside interrupts every sentence.

  1. Tell Vizard Agent there are two colours of text on screen.
  2. Say what you want back — timestamps, translation, or both.
  3. Ask for the two kinds separated rather than merged.

What do you need before you start?

The video and, if you know it, which colour means what. Vizard Agent can work it out by comparing the text against the speech, but confirming that green is commentary saves a pass and removes the main way this job goes wrong.

What do you type into Vizard Agent?

Name the colours explicitly. A request to "translate the subtitles" gets you one merged track, because to a transcriber every line of text on screen is a subtitle, and the distinction you care about is a visual one.

Prompt

Translate this video for me. Split it by timestamps, and keep the white text and the green text separate — I think the white is what the people are saying and the green is the narrator's note, but check. I need both, side by side, with the times.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on a short-form video carrying Chinese dialogue in white with green annotations above it. Almost all of the effort goes into reading small text accurately rather than into translating it, which is the opposite of what people expect from a translation job.

  1. Extracts frames to see how the text is laid out.
  2. Pulls a denser set of frames so every subtitle change is captured.
  3. Opens a full-size frame to fix the position and colour of each line.
  4. Zooms into the subtitle band at the moments there is dialogue.
  5. Separates the two text regions and enlarges each one, because small characters are misread at frame size.
  6. Runs automatic recognition on both colours to cross-check its own reading.
  7. Compares several timestamps to see what the green text changes with.
  8. Checks the green text at the end to confirm it is annotation rather than dialogue.
  9. Cross-checks with an analysis pass that separates the two by timestamp.
  10. Produces the timed translation with white and green kept apart.
  11. Repeats the whole process for the next video the user uploads.

Step five is the difference between a usable transcript and a plausible one. Dense Chinese characters at subtitle size are a few dozen pixels tall; read from the whole frame they come back as similar-looking characters that produce a fluent, confident and wrong translation.

Step eight is the check most people never make. If the green line is sometimes speech, the whole separation is invalid, so Vizard Agent tests the assumption at a point late in the video rather than carrying it from the first frame.

Step six is a cross-check, not a shortcut. Automatic recognition and reading the enlarged frames are two independent routes to the same text, and where they disagree is exactly where a human would have misread it too.

What does the result look like?

A timed document with two columns or two blocks per timestamp: what was said, and what the on-screen note added. The translation reads as dialogue rather than as a stream of interruptions, and you can see at a glance which parts of the meaning came from speech and which from the annotation.

When does this not work well?

Some text cannot be separated by colour. If the two kinds share a colour, if the annotation sometimes appears in the dialogue position, or if the video is heavily compressed so the colours bleed, the separation has no reliable signal to work from.

How do you fix a result that came back wrong?

Point at one timestamp and say what it should have said instead. Vizard Agent re-reads that moment at higher magnification rather than redoing the whole video, and a single correction of that kind often reveals a pattern it can then apply through the rest of the transcript.

How does Vizard Agent compare to doing it yourself?

By hand you pause, squint, type, and repeat for every line, in a language you may not read. The colours are the easy part for a person and the hard part is the volume; for a machine it is the reverse, which is why the enlargement steps exist.

By hand Vizard Agent
Reading Pause and squint Frames enlarged per region
Accuracy Whatever you could make out Two independent readings, cross-checked
The colour rule Assumed Tested late in the video
Timestamps Written down by hand Taken from the frames themselves
Output One mixed list Two tracks on one timeline

Common questions

Does it need to be two colours? No. Vizard Agent can split by position too, if the two kinds sit consistently apart.

Can it do three kinds? Yes, if each is visually distinct. Tell Vizard Agent what to look for.

What if I do not read the language? That is the usual case. Vizard Agent gives you the original and the translation.

Can I get a subtitle file? Yes. Ask Vizard Agent for SRT and the timings come with it.

Will it keep the original text? Yes, alongside the translation, so you can check anything that reads oddly.

Does it use the audio as well? Where there is audio, yes — it is how Vizard Agent confirms which colour is dialogue.

What about text that appears for half a second? Vizard Agent's dense frame pass is there to catch exactly those.

Can it burn new subtitles back in? Yes, though removing the old ones first is a separate job.

How accurate is the reading? Good when the text is legible enlarged; Vizard Agent flags what it is unsure of.

Can it handle several videos? Yes. Each one goes through the same process.

What if the annotation is a joke rather than a fact? It is kept as annotation. Vizard Agent separates by kind, not by tone.

Will the timings match the video exactly? Yes. Vizard Agent takes them from the frames where the text actually changes.

Can it tell me only the dialogue? Yes. Ask Vizard Agent for one track and the other is dropped.

Does it work on vertical video? Yes, and vertical is where this two-layer style is most common.