How to separate the two colours of subtitle burnt into a video
Tell Vizard Agent that the video carries two kinds of burnt-in text in two colours and that you need them kept apart. It reads each colour in its own region of the frame, works out which one is dialogue and which is annotation, and returns a timed translation with the two tracks separated.
What is the short version?
Short-form video from some markets carries two text layers at once: the spoken line in white along the bottom, and a green note explaining what is happening. Transcribe them together and you get a document where a narrator's aside interrupts every sentence.
- Tell Vizard Agent there are two colours of text on screen.
- Say what you want back — timestamps, translation, or both.
- Ask for the two kinds separated rather than merged.
What do you need before you start?
The video and, if you know it, which colour means what. Vizard Agent can work it out by comparing the text against the speech, but confirming that green is commentary saves a pass and removes the main way this job goes wrong.
- The video. With the text burnt in.
- The two colours. White and green, yellow and white.
- Which is dialogue. If you know.
- The target language. For the translation.
- What you want back. A document, a subtitle file, or both.
What do you type into Vizard Agent?
Name the colours explicitly. A request to "translate the subtitles" gets you one merged track, because to a transcriber every line of text on screen is a subtitle, and the distinction you care about is a visual one.
Prompt
Variants worth knowing:
- "Keep the white and the green separate." The whole job, in one line.
- "I think the white is speech, but check." Invites verification.
- "Side by side, with the times." The shape of the deliverable.
What does Vizard Agent actually do?
Here is the Vizard Agent sequence on a short-form video carrying Chinese dialogue in white with green annotations above it. Almost all of the effort goes into reading small text accurately rather than into translating it, which is the opposite of what people expect from a translation job.
- Extracts frames to see how the text is laid out.
- Pulls a denser set of frames so every subtitle change is captured.
- Opens a full-size frame to fix the position and colour of each line.
- Zooms into the subtitle band at the moments there is dialogue.
- Separates the two text regions and enlarges each one, because small characters are misread at frame size.
- Runs automatic recognition on both colours to cross-check its own reading.
- Compares several timestamps to see what the green text changes with.
- Checks the green text at the end to confirm it is annotation rather than dialogue.
- Cross-checks with an analysis pass that separates the two by timestamp.
- Produces the timed translation with white and green kept apart.
- Repeats the whole process for the next video the user uploads.
Step five is the difference between a usable transcript and a plausible one. Dense Chinese characters at subtitle size are a few dozen pixels tall; read from the whole frame they come back as similar-looking characters that produce a fluent, confident and wrong translation.
Step eight is the check most people never make. If the green line is sometimes speech, the whole separation is invalid, so Vizard Agent tests the assumption at a point late in the video rather than carrying it from the first frame.
Step six is a cross-check, not a shortcut. Automatic recognition and reading the enlarged frames are two independent routes to the same text, and where they disagree is exactly where a human would have misread it too.
What does the result look like?
A timed document with two columns or two blocks per timestamp: what was said, and what the on-screen note added. The translation reads as dialogue rather than as a stream of interruptions, and you can see at a glance which parts of the meaning came from speech and which from the annotation.
When does this not work well?
Some text cannot be separated by colour. If the two kinds share a colour, if the annotation sometimes appears in the dialogue position, or if the video is heavily compressed so the colours bleed, the separation has no reliable signal to work from.
- Both kinds the same colour. Position may still work, if it is consistent.
- Text that changes colour for emphasis. Colour stops meaning kind.
- Heavy compression. Thin coloured strokes wash out.
- Overlapping text. Two layers in the same band.
- Very small text on a busy background. Neither reading route is reliable.
How do you fix a result that came back wrong?
Point at one timestamp and say what it should have said instead. Vizard Agent re-reads that moment at higher magnification rather than redoing the whole video, and a single correction of that kind often reveals a pattern it can then apply through the rest of the transcript.
- "That line is in the wrong column." That timestamp is re-read by region.
- "This character is wrong." That frame is enlarged and checked again.
- "The green is speech here." The assumption is re-tested across the video.
- "I need it as a subtitle file." The same timings are exported as SRT.
How does Vizard Agent compare to doing it yourself?
By hand you pause, squint, type, and repeat for every line, in a language you may not read. The colours are the easy part for a person and the hard part is the volume; for a machine it is the reverse, which is why the enlargement steps exist.
| By hand | Vizard Agent | |
|---|---|---|
| Reading | Pause and squint | Frames enlarged per region |
| Accuracy | Whatever you could make out | Two independent readings, cross-checked |
| The colour rule | Assumed | Tested late in the video |
| Timestamps | Written down by hand | Taken from the frames themselves |
| Output | One mixed list | Two tracks on one timeline |
Common questions
Does it need to be two colours? No. Vizard Agent can split by position too, if the two kinds sit consistently apart.
Can it do three kinds? Yes, if each is visually distinct. Tell Vizard Agent what to look for.
What if I do not read the language? That is the usual case. Vizard Agent gives you the original and the translation.
Can I get a subtitle file? Yes. Ask Vizard Agent for SRT and the timings come with it.
Will it keep the original text? Yes, alongside the translation, so you can check anything that reads oddly.
Does it use the audio as well? Where there is audio, yes — it is how Vizard Agent confirms which colour is dialogue.
What about text that appears for half a second? Vizard Agent's dense frame pass is there to catch exactly those.
Can it burn new subtitles back in? Yes, though removing the old ones first is a separate job.
How accurate is the reading? Good when the text is legible enlarged; Vizard Agent flags what it is unsure of.
Can it handle several videos? Yes. Each one goes through the same process.
What if the annotation is a joke rather than a fact? It is kept as annotation. Vizard Agent separates by kind, not by tone.
Will the timings match the video exactly? Yes. Vizard Agent takes them from the frames where the text actually changes.
Can it tell me only the dialogue? Yes. Ask Vizard Agent for one track and the other is dropped.
Does it work on vertical video? Yes, and vertical is where this two-layer style is most common.