Vizard Agent

How to trust the burnt-in captions over the audio when a name is misheard

Last updated 2026-09-15 · 8 min read

Tell Vizard Agent that the video already has the words on screen and to read those instead. Speech recognition turns unfamiliar brand names into ordinary words that sound similar, and once that is in the transcript every graphic and every b-roll choice built from it is about something else entirely. Reading the burnt-in captions off the frames settles it.

What is the short version?

A transcript is a guess about sounds. A caption already burnt into the picture is a decision somebody made about the words themselves, and wherever the two disagree about a name, the one on screen is almost always the right one.

  1. Go to Vizard Agent with the video.
  2. Say the captions are burnt into the picture.
  3. Ask it to read those rather than transcribe the audio.

What do you need before you start?

The video with its captions visible, and the words you know it is getting wrong. Vizard Agent can read text off frames, but knowing which word is the problem tells it where to look and stops it re-transcribing the whole thing.

What do you type into Vizard Agent?

Point Vizard Agent at the disagreement rather than asking it for another transcription. Running speech recognition again produces exactly the same mishearing, because the audio has not changed and neither has the set of words the recogniser expects to find in it.

Prompt

The transcript has the brand name wrong — it hears an ordinary word that sounds similar. The correct name is already written in the captions burnt into the video. Read the text off the frames at those moments and use that, rather than transcribing the audio again.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the sequence on a real edit where b-roll had to match what was being said. The first pass illustrated a word the speaker never used, and the fix was to change where the words came from.

  1. Transcribes the speech to find the moments.
  2. Builds material against that first reading.
  3. Is told the name is wrong.
  4. Reads the captions directly off the video's frames.
  5. Takes the key moments from the on-screen text.
  6. Rebuilds the graphics against the corrected wording.
  7. Re-places the b-roll at the right moments.
  8. Checks each insert against what is actually on screen.

Step three is worth noticing as a habit rather than an error. When a user says a word is wrong, the useful response is not to try harder at listening; it is to ask whether a more reliable source for that word exists, and on a captioned video one always does.

Step four is the switch that fixes it. Text in a picture can be read rather than inferred, so a name that defeats speech recognition — an invented word, a handle, a product — is legible in a frame even when it is unrecoverable from the audio.

Step six is where the cost of the original mistake shows up. Everything built against a mis-heard word — a graphic, a search term, a caption line — has to be made again, which is why catching the wrong reading early is worth more than any amount of careful work afterwards.

Step eight is what keeps the correction honest. A graphic built on the right word can still land on the wrong moment, so the inserts get checked against the frames they sit on rather than against the transcript's timings alone.

What does the result look like?

Graphics and b-roll that match what the video actually says, with the brand name spelled the way the brand spells it. Vizard Agent tells you which moments it took from the on-screen text, so you can confirm the ones that matter.

When does this not work well?

This only helps where text exists. A video with no captions, captions that are themselves wrong, or a name that is spoken but never written leaves you back with the audio and a spelling you have to supply yourself.

How do you fix a result that came back wrong?

Give Vizard Agent the spelling directly rather than describing the sound. Because it is working from a word rather than from audio, handing over the correct spelling fixes every use of that name across the whole video at once rather than one graphic at a time.

How does Vizard Agent compare to doing it yourself?

By hand you would re-listen, hear the same thing, and eventually type the name in manually — after the graphics were already built around the wrong one. The on-screen text was there the whole time and nobody thought to read it.

By hand Vizard Agent
The source of truth The audio The text already in frame
When a word is wrong Listen again Read the frames instead
Brand names Mangled quietly Taken from the captions
Scope of the fix One graphic Every use of that word
Verification The transcript The frames the inserts sit on

Common questions

Why does transcription get names wrong? Because it fits sounds to words it expects, and invented names are not among them.

Can it read any on-screen text? Yes, if it is legible. Vizard Agent reads it off the frames.

What if there are no captions? Then tell Vizard Agent the spelling and it uses that.

Will it re-transcribe everything? No. It reads the moments in question rather than starting over.

What about handles and hashtags? Same problem, same fix. Vizard Agent reads those off the screen too.

Can it check the whole video? Yes. Ask Vizard Agent to compare the transcript against the on-screen text.

Does it fix the subtitles too? Yes, where Vizard Agent generated them from the same transcript.

What if the captions are stylised? Harder, and Vizard Agent will say when it cannot read them.

Will it correct the spelling everywhere? Yes. Vizard Agent applies it to graphics that were already built.

Can it get timings from the captions? Vizard Agent uses them to confirm placement, alongside the audio timings.

Does this apply to numbers? Yes. Vizard Agent treats percentages and prices the same way; they are misheard just as often.

What if the speaker says it wrong? Then the captions and the audio genuinely differ. Tell Vizard Agent which to use.

Is the audio ever more reliable? For ordinary words, yes, and Vizard Agent uses it. For names, the screen wins.

Why did the b-roll show the wrong thing? Because it was chosen for a word the speaker never said.