How to trust the burnt-in captions over the audio when a name is misheard
Tell Vizard Agent that the video already has the words on screen and to read those instead. Speech recognition turns unfamiliar brand names into ordinary words that sound similar, and once that is in the transcript every graphic and every b-roll choice built from it is about something else entirely. Reading the burnt-in captions off the frames settles it.
What is the short version?
A transcript is a guess about sounds. A caption already burnt into the picture is a decision somebody made about the words themselves, and wherever the two disagree about a name, the one on screen is almost always the right one.
- Go to Vizard Agent with the video.
- Say the captions are burnt into the picture.
- Ask it to read those rather than transcribe the audio.
What do you need before you start?
The video with its captions visible, and the words you know it is getting wrong. Vizard Agent can read text off frames, but knowing which word is the problem tells it where to look and stops it re-transcribing the whole thing.
- The video. With the captions in the picture.
- The word it got wrong. And what it should be.
- Whether captions cover the whole video. Or just parts.
- What the transcript is for. Graphics, b-roll, subtitles.
- Any other proper nouns. They fail the same way.
What do you type into Vizard Agent?
Point Vizard Agent at the disagreement rather than asking it for another transcription. Running speech recognition again produces exactly the same mishearing, because the audio has not changed and neither has the set of words the recogniser expects to find in it.
Prompt
Variants worth knowing:
- "Already written in the captions." Names the better source.
- "Read the text off the frames." The method.
- "Rather than transcribing again." Stops the loop.
What does Vizard Agent actually do?
Here is the sequence on a real edit where b-roll had to match what was being said. The first pass illustrated a word the speaker never used, and the fix was to change where the words came from.
- Transcribes the speech to find the moments.
- Builds material against that first reading.
- Is told the name is wrong.
- Reads the captions directly off the video's frames.
- Takes the key moments from the on-screen text.
- Rebuilds the graphics against the corrected wording.
- Re-places the b-roll at the right moments.
- Checks each insert against what is actually on screen.
Step three is worth noticing as a habit rather than an error. When a user says a word is wrong, the useful response is not to try harder at listening; it is to ask whether a more reliable source for that word exists, and on a captioned video one always does.
Step four is the switch that fixes it. Text in a picture can be read rather than inferred, so a name that defeats speech recognition — an invented word, a handle, a product — is legible in a frame even when it is unrecoverable from the audio.
Step six is where the cost of the original mistake shows up. Everything built against a mis-heard word — a graphic, a search term, a caption line — has to be made again, which is why catching the wrong reading early is worth more than any amount of careful work afterwards.
Step eight is what keeps the correction honest. A graphic built on the right word can still land on the wrong moment, so the inserts get checked against the frames they sit on rather than against the transcript's timings alone.
What does the result look like?
Graphics and b-roll that match what the video actually says, with the brand name spelled the way the brand spells it. Vizard Agent tells you which moments it took from the on-screen text, so you can confirm the ones that matter.
When does this not work well?
This only helps where text exists. A video with no captions, captions that are themselves wrong, or a name that is spoken but never written leaves you back with the audio and a spelling you have to supply yourself.
- No captions on screen. Nothing to read.
- Captions that are wrong too. They inherit the same error.
- Names spoken but never written. Tell it the spelling.
- Stylised or animated text. Harder to read reliably.
- Fast captions. A word may never sit still on a frame.
How do you fix a result that came back wrong?
Give Vizard Agent the spelling directly rather than describing the sound. Because it is working from a word rather than from audio, handing over the correct spelling fixes every use of that name across the whole video at once rather than one graphic at a time.
- "It is spelled like this." Applied everywhere it appears.
- "The captions are wrong there too." Your spelling wins.
- "Wrong moment, right word." The placement is re-timed.
- "There is no caption for it." Supply the word yourself.
How does Vizard Agent compare to doing it yourself?
By hand you would re-listen, hear the same thing, and eventually type the name in manually — after the graphics were already built around the wrong one. The on-screen text was there the whole time and nobody thought to read it.
| By hand | Vizard Agent | |
|---|---|---|
| The source of truth | The audio | The text already in frame |
| When a word is wrong | Listen again | Read the frames instead |
| Brand names | Mangled quietly | Taken from the captions |
| Scope of the fix | One graphic | Every use of that word |
| Verification | The transcript | The frames the inserts sit on |
Common questions
Why does transcription get names wrong? Because it fits sounds to words it expects, and invented names are not among them.
Can it read any on-screen text? Yes, if it is legible. Vizard Agent reads it off the frames.
What if there are no captions? Then tell Vizard Agent the spelling and it uses that.
Will it re-transcribe everything? No. It reads the moments in question rather than starting over.
What about handles and hashtags? Same problem, same fix. Vizard Agent reads those off the screen too.
Can it check the whole video? Yes. Ask Vizard Agent to compare the transcript against the on-screen text.
Does it fix the subtitles too? Yes, where Vizard Agent generated them from the same transcript.
What if the captions are stylised? Harder, and Vizard Agent will say when it cannot read them.
Will it correct the spelling everywhere? Yes. Vizard Agent applies it to graphics that were already built.
Can it get timings from the captions? Vizard Agent uses them to confirm placement, alongside the audio timings.
Does this apply to numbers? Yes. Vizard Agent treats percentages and prices the same way; they are misheard just as often.
What if the speaker says it wrong? Then the captions and the audio genuinely differ. Tell Vizard Agent which to use.
Is the audio ever more reliable? For ordinary words, yes, and Vizard Agent uses it. For names, the screen wins.
Why did the b-roll show the wrong thing? Because it was chosen for a word the speaker never said.