Vizard Agent

How to caption the sounds and the speakers, not just the words

Last updated 2026-09-17 · 8 min read

Ask Vizard Agent for speaker-labelled captions with sound cues, not just a transcript on screen. It works out who is speaking by sampling the mouths in the picture at each line rather than trusting the audio separation alone, and adds cues for the sounds that carry meaning — a knock, a car pulling away — while leaving the ambient ones out.

What is the short version?

An ordinary caption track is a transcript with timings attached to it. A viewer who cannot hear the video gets all of the words and none of the other information the soundtrack was carrying: which person said them, and what just happened off screen.

  1. Ask Vizard Agent for speaker labels and sound cues.
  2. Ask it to verify each speaker against the picture.
  3. Say which sounds matter and which are just ambience.

What do you need before you start?

The video and the names of the people in it. Vizard Agent can establish that two different people are speaking; which of them is Maria and which is the neighbour is something only somebody who has actually watched the film can supply.

What do you type into Vizard Agent?

Ask Vizard Agent for the two additions by name. A phrase like "captions for the deaf" means slightly different things in different places, and spelling out "speaker labels and sound cues" removes all of that ambiguity in a single sentence.

Prompt

Caption this for viewers who cannot hear it. I want speaker labels on every line and cues for the sounds that matter — the knock at the door, the car leaving — but not for general ambience. Check who is speaking against the picture rather than guessing from the audio, and keep the labels readable.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on a real short film captioned for accessibility. The interesting part is step four, where the attribution is checked against the mouths visible on screen rather than simply accepted from the audio analysis.

  1. Listens through for dialogue, music and sound-effect events.
  2. Builds timed captions from the dialogue transcript.
  3. Inspects the transcript structure for its speaker information.
  4. Samples the speaker's visible mouth movements at each dialogue cue.
  5. Reviews the speaker frames across every section of the video.
  6. Builds speaker-labelled dialogue and sound-cue caption data.
  7. Tests the labelled captions and cues on the opening section.
  8. Checks the labels and cues for readability at real size.
  9. Repositions the caption block into the visible safe area.
  10. Renders the full video with labels and sound cues in place.

Step four is the one that turns a guess into an attribution. Audio separation tells you there are two voices; looking at whose mouth is moving at that moment tells you which one is which, and getting that wrong is worse than having no labels at all.

Step eight matters more than it does for plain captions. A label and a cue both take room on the line, so a caption that was comfortable becomes two crowded lines, and the readability has to be rechecked after they are added.

Step one is a listening pass rather than a transcription pass. Sound cues do not appear in a transcript, so the events have to be found by listening to the track for them.

What does the result look like?

Captions that carry the soundtrack rather than only the script. Each line says who is speaking, the sounds that actually move the story along are cued while the ambient ones are left out, and Vizard Agent has kept the block sitting where it can still be read comfortably.

When does this not work well?

Attribution needs faces to check against. An off-screen voice, a crowd scene, or a conversation shot from behind gives Vizard Agent nothing to verify against, and those lines have to be labelled from context or left unlabelled rather than guessed at.

How do you fix a result that came back wrong?

Say whether it is a label that is wrong or the sound cues. The two come from entirely different passes — one reads the picture, the other listens to the soundtrack — so Vizard Agent adjusts them separately rather than rebuilding the whole track.

How does Vizard Agent compare to doing it yourself?

By hand you caption the dialogue, because that is what the transcript gives you, and adding who-said-what means watching the whole thing again with a notepad. Sound cues mean a third pass, listening only for events, so most captions never get either.

By hand Vizard Agent
What the track contains The words Words, speakers and sounds
Attribution From the transcript's guess Checked against the mouths
Sound cues A separate listening pass Found in the same review
Readability after labels Rarely rechecked Re-measured at real size
Off-screen voices Guessed Flagged rather than invented

Common questions

Is this the same as subtitles? No. Subtitles carry the words; this carries who said them and what was heard.

How should speakers be named? By name if the viewer knows them, by role if not. Tell Vizard Agent which.

Which sounds should be cued? The ones that carry meaning. Vizard Agent leaves ambience out unless you ask.

What about music? It can be cued — the title, or a description of its mood. Say which you prefer.

Will the labels fit? Vizard Agent rechecks readability after adding them, because they take room.

What if the speaker is off screen? Vizard Agent flags it rather than guessing, and you decide how to label it.

Can it produce a sidecar file? Yes. Vizard Agent delivers one as well as, or instead of, burnt-in captions.

Does this meet a particular standard? Tell Vizard Agent which one and it works to that.

What about two people talking at once? Hard to label cleanly, so Vizard Agent shows you the options rather than picking.

Can it use colours per speaker? Yes. Vizard Agent can use that convention, which saves room on the line.

Will it slow the captions down? Labels add length, so some lines need more time. Vizard Agent adjusts.

Does it work on a documentary? Yes, and interviews benefit most from labels.

Can I get a plain version too? Yes. Ask Vizard Agent for both tracks.

How do I check it is right? Watch it with the sound off. That is the experience it is built for.