Vizard Agent

How to stop a name being spelled wrong in every clip you publish

Last updated 2026-09-22 · 8 min read

Give Vizard Agent the exact spellings before it cuts anything. Speech recognition does not know how a surname is spelled, so it produces a plausible version — and in a batch of clips that plausible version appears on every name label and every caption at once.

What is the short version?

You are cutting one episode into six clips. Every one of them carries two names and a show title on screen, so a single wrong spelling from Vizard Agent or anyone else becomes six public mistakes rather than one.

  1. Give Vizard Agent the names and the show title, spelled exactly.
  2. Say who is who — which speaker sits on which side.
  3. Check the first export before the rest are rendered.

What do you need before you start?

The episode and the spellings. Vizard Agent reads the transcript and the frames, but a surname with two possible spellings is a coin toss without you, and the show's own title is often styled in a way no transcript would guess.

What do you type into Vizard Agent?

Put the spellings in the first message as a list. It takes one line and it is the difference between six correct clips and six corrections, and it also gives Vizard Agent something to check the captions against as it goes.

Prompt

Cut my podcast into clips of 10 to 30 seconds. Make sure our names and the show title are correct: Melissa Cunningham, Sasha Eburne, The Sasha Eburne Show. Label each speaker properly, keep the captions clear of our faces, and show me one finished clip before you render the rest.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on an episode cut into six clips with two speakers in shot. All of the identification work happens before any clip is built, because a name label ends up attached to every single one of them.

  1. Reviews the episode for strong ten-to-thirty-second clips, using the exact names supplied.
  2. Checks the conversation, the camera layout and the introductions before choosing anything.
  3. Reads the transcript against the actual frames so each clip credits the right speaker.
  4. Checks how both speakers are framed and how any existing name labels appear.
  5. Confirms which speaker is on which side of the frame.
  6. Compares short self-contained passages for the strongest candidates.
  7. Separates complete advice from answers that need too much setup.
  8. Narrows the batch across the middle and closing sections.
  9. Locks six cut points and checks a closer frame before building the layout.
  10. Checks face positions at full resolution so the portrait layout clears both faces.
  11. Builds the shared titles and name labels and measures the clips' sound.
  12. Adds word-timed captions and renders one layout check before exporting the batch.
  13. Reads the render error when the first export fails, corrects an audio-format mismatch, and checks the caption size.
  14. Reviews the portrait crop, the exact spellings and the caption placement on that first export.

Step five is a small thing that decides whether the labels mean anything. Getting the names right and then putting them on the wrong sides is a worse mistake than a misspelling, because it looks deliberate.

Step twelve is the batch discipline. Rendering one clip and looking at it costs a minute; rendering six and finding the caption size wrong costs six times that plus the re-export.

Step three matters because a transcript alone cannot attribute lines when two people are in shot. Reading it against the frames is what keeps each clip credited to whoever is actually speaking in it.

What does the result look like?

Six clips with the names spelled the way their owners spell them, the right name against the right face, a consistent layout, captions clear of both faces, and one of them checked in full before the rest were made.

When does this not work well?

Some episodes cannot be labelled reliably at all. If the speakers are never both in shot, if they swap seats partway through, or if a guest is never introduced by name, then the attribution has to come from you rather than from anything Vizard Agent can read.

How do you fix a result that came back wrong?

Correct the spelling once and it propagates through the batch. Vizard Agent holds the names as data for the whole job rather than as text typed into each clip separately, so one correction fixes every clip that carries that name.

How does Vizard Agent compare to doing it yourself?

By hand you type the names into the first clip's template, duplicate it five times, and discover at clip four that you spelled the guest's surname the common way rather than her way. Then you fix five files.

By hand Vizard Agent
Spellings Typed per clip Held as data for the batch
Attribution From memory Transcript checked against the frames
Layout Copied from the first clip Built from measured face positions
Errors Repeated across the batch Corrected once, propagated
Checking After exporting them all One clip first, then the rest

Common questions

Why does speech recognition get names wrong? It hears sounds, not spellings — which is why Vizard Agent asks for them.

Can it label speakers automatically? Vizard Agent attributes by voice and position, and you confirm which is which.

What if my show title has odd punctuation? Give Vizard Agent the exact form and it uses that everywhere.

Will it caption the clips too? Yes, word-timed, with the corrected spellings included.

Can it use our brand colours? Yes, in the labels and the layout across the batch.

What if one clip should have a different label? Say which and Vizard Agent changes that one alone.

Can it credit a guest's company as well? Yes — name, role and company all sit in the same label.

How many clips from one episode? As many as stand alone; Vizard Agent will say what it found.

Will the clips look consistent? Yes. Vizard Agent builds one layout and applies it across the batch.

Can I check one before the rest? Yes. Ask Vizard Agent for one finished clip first — it is worth doing every time.

What about names in the captions themselves? They are corrected there too, from the same list.

Does it work for three or more speakers? Yes, though attribution gets harder as more people overlap.

Can it produce horizontal versions? Yes, from the same cut points and the same labels.

What if we rebrand the show? Give Vizard Agent the new title and the next batch carries it.