How to give each character in a dub their own voice
Give Vizard Agent a cast list rather than a language, and say that characters of the same gender need voices you can tell apart. It generates each character's lines into the existing timeline's windows so nothing drifts, mixes them into one dialogue track, then re-derives the captions from that mixed track rather than from the version they were timed to.
What is the short version?
Dubbing is usually set up as a single job — this language into that one — so a single voice ends up reading the whole thing. In any scene with two people talking to each other, that turns a conversation into a monologue delivered in two halves.
- Tell Vizard Agent who the characters are.
- Say which voices must be distinguishable from each other.
- Ask for the captions rebuilt from the new mixed track.
What do you need before you start?
The video and a list of who speaks. Vizard Agent can hear that there are several speakers, but which of them is the husband and which the shopkeeper is something only somebody who has watched it can say.
- The video. With its original dialogue.
- The cast. Who speaks, and roughly who they are.
- Gender and age per character. It drives the casting.
- Who must sound different. Two men, two women.
- Whether captions are wanted. They need rebuilding after.
What do you type into Vizard Agent?
Name the characters and their relationships. "Use male and female voices" is a rule about gender; "the husband and the shop owner should not sound like the same man" is a rule about the scene, and only the second one produces a usable dub.
Prompt
Variants worth knowing:
- "Not with one narrator." Rules out the default.
- "Clearly different men." The constraint people forget.
- "Rebuild the subtitles from the finished track." Otherwise they drift.
What does Vizard Agent actually do?
Here is the Vizard Agent sequence on a real dub that had to be rebuilt as a multi-voice track. The last third of it is entirely about keeping the captions honest, because every single change to the dialogue invalidates the timings they were originally built from.
- Finds two distinct voices for the two male characters.
- Generates each character's dialogue separately.
- Refines each line so it fits its scene naturally.
- Mixes the character voices into the existing dub.
- Generates replacements aligned to the existing timeline windows.
- Rebuilds the mixed track using the correct timing windows.
- Transcribes the mixed track rather than reusing the old transcript.
- Builds captions from that transcription.
- Applies strict gender assignment across the whole dialogue track.
- Inspects any shot where it is unclear who is speaking.
Step five is what keeps a multi-voice dub in sync. Each new line has a window it must fit, and a line generated freely and dropped in afterwards pushes everything downstream of it out of place.
Step seven is the step that gets skipped. Captions built before the voices changed describe a soundtrack that no longer exists, and they will be close enough to look plausible and wrong everywhere.
Step ten is where honesty beats guessing. If a shot does not show who is talking, assigning a voice is a coin toss, and it is worth resolving before the whole track is rebuilt around the assumption.
What does the result look like?
A scene you could follow with your eyes shut. Each character sounds like a different person, every line still lands inside its original window, and Vizard Agent derived the captions from the track that actually plays rather than from the one it replaced.
When does this not work well?
Some dialogue simply cannot be separated. Overlapping speech, crowd scenes and any passage where the original mix has two people sharing one track will not divide cleanly into per-character lines, and Vizard Agent will say so rather than force it.
- Overlapping dialogue. Two people, one window.
- Crowd or background chatter. Too many voices to cast.
- Characters you never see. No way to attribute the lines.
- Very short lines. A voice needs a phrase to establish itself.
- Strict lip sync. Different voices phrase at different lengths.
How do you fix a result that came back wrong?
Name the character rather than the timecode. Vizard Agent holds the whole dub as a set of per-character lines, so a note like "the shop owner sounds too young" recasts that one character everywhere he speaks without disturbing anybody else in the scene.
- "They sound like the same man." A more distinguishable voice is cast.
- "The shop owner sounds too young." That character is recast throughout.
- "The captions are off." They are rebuilt from the current mixed track.
- "Wrong voice on that line." The shot is inspected and the line reassigned.
How does Vizard Agent compare to doing it yourself?
By hand you dub it, hear that everyone sounds the same, and re-record the second character. Then the new line is a different length, the mix drifts, and the subtitles you had already approved no longer match anything.
| By hand | Vizard Agent | |
|---|---|---|
| The brief | A target language | A cast list |
| Same-gender characters | One voice for both | Voices chosen to differ |
| New lines | Recorded, then fitted | Generated into the window |
| Captions | Reused from before | Rebuilt from the mixed track |
| Unclear speakers | Guessed | Inspected, then assigned |
Common questions
Why does it default to one voice? Because a dub is usually briefed as a language, not as a cast. Give Vizard Agent the cast.
How many voices can it handle? As many as you list, though Vizard Agent struggles to characterise very short lines.
Can two men sound properly different? Yes, if you ask. Vizard Agent chooses for contrast rather than just gender.
Will the lines still fit? Yes. Vizard Agent generates each into its existing timing window.
Do the captions change? They must. Vizard Agent rebuilds them from the new dialogue track.
What about overlapping speech? It cannot be split cleanly. Vizard Agent will flag those passages.
Can it match the original actors' voices? Sometimes, with samples. Otherwise it casts for character rather than likeness.
What if I do not know who says what? Vizard Agent can transcribe with speaker labels first, then you name them.
Does lip sync still work? Loosely. Different voices phrase differently; Vizard Agent reports the fit.
Can I change one character later? Yes, and only that character's lines are regenerated.
What about a narrator as well? List them as a character and Vizard Agent casts them the most neutral voice.
Will the mix need rebalancing? Yes, if the voices differ in level. Vizard Agent measures them.
Can it keep the original audio underneath? Yes. Ask Vizard Agent to hold it at a low level if you want the original performance audible.
How do I check it worked? Listen to the Vizard Agent dub with the picture off. If you can still tell who is who, it worked.