How to transcribe a video into a dialogue script with the speaker names
Ask Vizard Agent for the dialogue laid out as a script rather than as a transcript. It extracts the spoken words with timings, pulls detailed frames of each character to see who is moving and speaking at each line, matches every sentence to a speaker, and writes the result with names against the lines — translated or in the original, as you prefer.
What is the short version?
Transcription gives you the words and nothing else. In a conversation the words are only half the information — the other half is who said each of them, and that turns out to be a question about the picture rather than about the audio.
- Go to Vizard Agent with the video.
- Ask for a script with speaker names, not a transcript.
- Say which language you want it in.
What do you need before you start?
The video and a decision about names and language. Vizard Agent can identify the characters from the footage and label them descriptively, but if they have names you know, giving them saves a round and makes the script readable straight away.
- The video. With the faces visible if possible.
- The characters' names. Or say to describe them.
- Which language. Original, translated, or both.
- Whether other languages come out. Subtitles, on-screen text.
- What the script is for. Reading, dubbing, or a rewrite.
What do you type into Vizard Agent?
Ask Vizard Agent for a script and state the language rule in the same instruction. A transcript and a script are two different deliverables, and a video with two languages in it will otherwise come back with both of them tangled together on the page.
Prompt
Variants worth knowing:
- "Laid out as a script." Not a paragraph of text.
- "Labelled by which character." The part that takes work.
- "One language only." Stops a bilingual mess.
What does Vizard Agent actually do?
Here is the sequence Vizard Agent worked through on a real two-character animated clip. The transcription itself is quick; almost all of the effort goes into deciding who is speaking on each line, and that is settled by looking rather than by listening.
- Samples frames to see what is in the video.
- Extracts the spoken words with timings.
- Reviews the characters who appear.
- Pulls a detailed preview of their movement.
- Inspects those frames to see who is speaking.
- Matches the speaker across the first half.
- Matches the speaker across the second half.
- Writes the dialogue file with names against each line.
- Produces a single-language version if you asked for one.
Step two is worth separating from the rest. Getting the words down with their timings is the cheap, reliable half of the job, and having it finished before attribution starts means a mistaken speaker label never costs you the transcription work underneath it.
Step four is the technique that makes this work on animation and on footage where voices are similar. Mouth movement, gesture and who is facing the camera are visible in a frame strip, and a strip at a fine enough interval settles a line that audio alone leaves ambiguous.
Step six and step seven are separate on purpose. Speaker matching drifts — one wrong attribution propagates until something obvious contradicts it — so working through the video in halves means a mistake is contained rather than inherited to the end.
Step nine is worth asking for explicitly. A video with two languages in it produces a transcript with two languages in it, and a script that mixes them is unusable for the thing you almost certainly want it for.
What does the result look like?
A readable script: the character's name, then their line, in order, in a single language throughout. Vizard Agent delivers it as a file rather than as text in the chat, which matters a great deal once it runs to several pages.
When does this not work well?
Attribution needs something visible to go on, and Vizard Agent will say when it has nothing. Off-screen speakers, crowd scenes, overlapping dialogue and heavy stylisation all make it hard to tell who is talking, and a guess in a script is considerably worse than an honest gap.
- Off-screen voices. Nothing visible to attribute to.
- Overlapping dialogue. Two speakers, one line.
- Crowds. Too many candidates.
- Narration over action. The narrator is nobody on screen.
- Poor audio. The words come first, and they may not.
How do you fix a result that came back wrong?
Point at a specific line and say who actually said it. Because Vizard Agent worked through the video in sections rather than line by line, one corrected attribution usually repairs a whole run of lines rather than only the single one you named.
- "That line is the other character." Corrected, and the run re-checked.
- "Use their real names." Applied throughout.
- "Keep the original language too." A second version.
- "There is a narrator as well." Labelled separately.
How does Vizard Agent compare to doing it yourself?
By hand you transcribe, then scrub back through the video attributing lines one at a time, which is slow and gets less accurate as attention drops. Most people give up and deliver an unlabelled transcript with an apology.
| By hand | Vizard Agent | |
|---|---|---|
| The words | Transcribed | Transcribed with timings |
| Who said it | Scrubbed for, line by line | Read off frames of the speakers |
| Drift | Compounds down the page | Contained by working in halves |
| Two languages | Tangled together | Separated on request |
| The deliverable | A transcript | A script, as a file |
Common questions
Why not use the audio to tell speakers apart? It helps, but Vizard Agent uses the picture to settle the cases audio cannot.
Can it name the characters? Yes, if you give the names. Otherwise Vizard Agent describes them.
Does it work on animation? Yes, and there the picture is often the only reliable signal.
What about a narrator? Vizard Agent labels them separately from the on-screen characters.
Can I have both languages? Yes, as two files. Vizard Agent avoids one mixed file, which is rarely useful.
Will it include timings? If you want them. Tell Vizard Agent whether the script is for reading or for editing.
What if two people talk at once? Vizard Agent marks it rather than guessing at a split.
Can it handle more than two characters? Yes, though Vizard Agent finds attribution harder as the count rises.
Does it deliver a file? Yes. Vizard Agent writes a file, which matters once the script runs long.
What about on-screen text? Ask Vizard Agent to include it; it is not dialogue but it often matters.
Can it rewrite the script afterwards? Yes, though that is a separate instruction from transcribing it.
Will it keep the original wording? Yes. Vizard Agent changes it only if you ask for a translation or a rewrite.
Can it do this for several videos? Yes, and the same character labels carry across them.
Why is attribution the slow part? Because it is a question about the picture, and the picture has to be looked at.