Vizard Agent

How to make an audio description track for blind viewers

Last updated 2026-09-28 · 8 min read

Ask Vizard Agent to describe only what is visible — the actions in frame and the changes of shot — and to fit that description into the gaps in the existing soundtrack. Every phrase gets measured against the space available, and the on-screen titles get read aloud, because that text is information the audience has no other way to reach.

What is the short version?

You have a finished video and you want it usable by someone who cannot see it. That is not a voiceover laid on top; it is a commentary written to fit between the sounds already there. Vizard Agent works from the gaps outwards.

  1. Tell Vizard Agent to describe the picture only, for blind viewers.
  2. Let it measure each phrase against the gap it has to sit in.
  3. Ask for the voice track on its own as well as mixed.

What do you need before you start?

The video, and a decision about register. A plain factual description and a more interpretive one are both legitimate choices that suit different material — a documentary usually wants the first, while an exhibition film or an art piece sometimes wants the second, and Vizard Agent can produce either.

Think about the voice. It will be heard over the whole piece alongside the original sound, so a calm, unhurried delivery matters more than character. In that session the voice was changed once to something warmer and more announcer-like after the first version was heard in place.

What do you type into Vizard Agent?

Tell Vizard Agent what to describe and, just as importantly, what to leave alone. The commonest mistake in this format is a description that narrates the plot or repeats back something the dialogue has already said perfectly clearly.

Make an audio description for this video. Describe only the visual track — the actions in frame and the changes of shot — for blind viewers.

The word "only" is the whole instruction. Everything audible is already reaching the listener, so the description's job is strictly the part that is not.

What does Vizard Agent actually do?

It maps the picture and the silence before writing anything. A shot-change map across the whole piece establishes where the visual information changes, and the original audio track is pulled down separately so the gaps in it can be measured — those gaps are the only places a description can go.

In that session the work ran:

Then the mix: the description's loudness and the original audio's loudness were measured at the same passage, the original was set underneath, and the balance was checked and adjusted again after listening.

What does the result look like?

Three deliverables rather than one: the timed description as text, the mixed track with the original sound under the commentary, and the voice track on its own. The separate voice track matters — it lets the description be used in a player that mixes it, or re-balanced later without regenerating anything.

The on-screen text is spoken. Opening titles, information cards, technical credits and the exhibition or programme name were all read off the frames, with one disputed word resolved by enlarging it — because a title nobody reads aloud is simply lost.

When does this not work well?

When there are no gaps. A densely narrated piece with wall-to-wall dialogue leaves nowhere to put a description, and the real options are an extended version that pauses the picture or accepting that some visual information cannot be conveyed. Vizard Agent will tell you which passages have no room.

Fast cutting is the other hard case. A sequence that changes shot every second cannot be described shot by shot at any speaking rate, so the description has to summarise the sequence instead — which is a judgement about what matters.

And it is not a compliance sign-off. Broadcast and public-sector accessibility standards carry specific requirements about how much must be covered, how it must be voiced and what format it must be delivered in, and meeting one of those standards is a separate exercise from producing a description that is genuinely good to listen to.

How do you fix a result that came back wrong?

If a description phrase collides with the dialogue, give Vizard Agent the timestamp. It shortens that phrase rather than compressing the speech to fit, because a description delivered too fast to follow is worse than a shorter one that leaves something out.

If two description phrases run into each other, say so. That is a distinct defect from colliding with the original audio and it gets checked separately — the fix is a gap between them, not a faster read.

If the register is wrong, ask for the other version. A factual pass and an interpretive pass can be produced from the same shot map, and hearing both against the picture is the only reliable way to choose.

How does Vizard Agent compare to doing it yourself?

Writing a description by hand means watching with a stopwatch, and the constraint that bites is not the writing — it is that the good sentence you wrote is four words too long for the gap. Most of the effort goes into trimming to fit, repeatedly.

Measuring every phrase against its gap up front turns that into arithmetic. The other thing worth copying is the delivery format: text, mixed track and isolated voice, so the description can be re-balanced or re-used without going back to the beginning.

Common questions

Should it describe the dialogue? No. Only what cannot be heard. Vizard Agent describes the picture alone.

What about on-screen text? It gets read aloud. Vizard Agent reads the titles and credits off the frames.

How loud should the description be? Above the original, which sits underneath it. Vizard Agent measures both rather than setting either by ear.

Can I get the text as well? Yes. Vizard Agent delivers it timed with timecodes alongside the audio.

Why a separate voice track? So it can be re-balanced or mixed in a player without asking Vizard Agent to regenerate anything.

What if there is no gap? Vizard Agent says which passages have no room rather than talking over the dialogue.

Can it use one consistent voice? Yes, and it should. Vizard Agent generates every fragment with the same voice.

Can the voice be changed later? Yes. Vizard Agent reuses the timed script with a different voice.

Is a factual or interpretive style better? Depends on the material. Vizard Agent can produce both from the same map.

Does it describe every shot change? Where there is room for it. Vizard Agent summarises fast sequences instead.

Will it meet a broadcast standard? Not automatically. Treat that as a separate specification.

Can it work from a link? Yes, if the video is reachable. Vizard Agent downloads it to work frame by frame.

What about a foreign-language original? The description can be in any language you ask for.

How long does it take? Most of the time goes on the shot map and the phrase fitting, not the voice.