How to direct a voiceover to change tone across the piece
Write the performance direction into the script section by section and let Vizard Agent generate each section separately in the same voice. A single long generation settles into one register and stays there; the arc you want — reflective, then analytical, then assertive — comes from generating the stretches apart and joining them.
What is the short version?
You have a script that should open quietly, sharpen in the middle and lift at the end. Handed over as one block it comes back level throughout, because nothing told the read where to turn. Vizard Agent cuts the script at the turns instead.
- Mark the tone for each stretch in the script you send Vizard Agent.
- Let it split at the idea transitions rather than at equal lengths.
- Have the shots and titles timed to those same joins.
What do you need before you start?
The script, with the direction written into it. A bracketed note before a section — calm and reflective, slower, with short pauses at the commas — is a complete instruction, and it is far more useful than a general description of the voice you want.
Decide the voice separately from the performance. "An experienced mentor, warm, grounded, not a radio announcer" describes the instrument; the arc describes what it plays. Those are two decisions and conflating them is why voiceovers come back sounding generically enthusiastic.
What do you type into Vizard Agent?
Put the voice profile at the top and the per-section direction inline, next to the words it applies to. What you hand Vizard Agent will look like an over-annotated script, and that is exactly what it should look like — every bracketed note is a decision it would otherwise have to guess at.
Warm, articulate, sounds like a mentor — not a radio announcer. Opening: calm and reflective, slow, short pauses. Middle: sharper, analytical. Close: assertive and encouraging.
The negative instruction matters as much as the positive ones. Generated narration drifts towards announcer by default, and naming that as the thing to avoid does more than any amount of describing what you do want.
What does Vizard Agent actually do?
It tests the voice before it commits to the script, then splits. A short line is generated first just to measure how fast that voice actually speaks, because the pace determines how much script fits each shot — and that has to be known before anything is apportioned.
In that session the sequence ran:
- Chose one voice and generated a single short sentence with it to measure its speed.
- Split the script into seven connected segments, cut at the points where the thinking turns rather than at equal durations.
- Generated all seven in parallel in the same voice, each with its own pacing direction.
- Joined them into one track and recorded where each section begins.
- Used those start points to place the titles and the shots, so the picture turns where the voice turns.
- Measured the narration and the music separately before mixing, so the words stayed clearly above the bed.
Splitting at the idea transitions is what makes the rest work. If the segments were equal lengths, the joins would fall mid-thought and the changes of register would sound like edits.
What does the result look like?
Narration that actually moves through the piece. The opening is unhurried, the middle is tighter and faster, the close pushes — and because Vizard Agent sits the visual beats on the same boundaries, each shift in register reads as intentional rather than as a splice between two takes.
The section starts become the edit's skeleton. Once you know where segment four begins you know where the fourth title card goes, and that alignment is free — it comes out of the same list rather than being matched up by hand afterwards.
When does this not work well?
When the script has no turns in it. A list of five tips with no argument between them gives nothing to change register for, and a performance arc laid over it sounds arbitrary. Vizard Agent will say when the script is flat rather than acting anyway.
Joins can be audible if the direction jumps too far. Calm straight into assertive in one cut is a different person; the fix is a middle segment, which is usually what the script wanted anyway.
And a very short piece has no room for this. Under about thirty seconds there is one register, and choosing the right one matters more than changing it.
How do you fix a result that came back wrong?
Name the section and the direction you wanted from it. "The third one is still too flat, push it harder" is immediately actionable, and Vizard Agent regenerates that one segment in the same voice rather than re-reading the entire script from the top.
If a join is audible, say where. That is usually two adjacent directions being too far apart, and the fix is either an intermediate segment or a softer boundary rather than a different voice.
If the whole thing sounds like an announcer, say that plainly. It is the default failure and it responds better to a negative instruction than to more adjectives about warmth.
How does Vizard Agent compare to doing it yourself?
A human narrator does this without being asked, because they can read the script and hear where it turns. That is exactly the thing a single generated read cannot do — it has no view of the argument, only of the sentence it is currently saying.
Segmenting is how you give it that view. The work is deciding where the turns are, which is a reading job you would do anyway before directing anyone, and the payoff is that the picture edit inherits those same boundaries for nothing.
Common questions
How many segments is right? As many as the script has turns. Seven suited a five-minute piece in this session.
Will it sound like one voice? Yes. Vizard Agent uses the same voice for every segment.
Why test a short line first? To measure the actual speaking pace, so Vizard Agent can apportion the script to the shots.
Can I write the direction in brackets? Yes, that is the clearest way. Vizard Agent reads it as direction rather than script.
What if a segment is too fast? Say which one and Vizard Agent regenerates only that segment.
Can the pauses be directed too? Yes. Short pauses after a key line are worth asking Vizard Agent for by name.
Does this work in any language? Yes. Vizard Agent measures the pace in the language of the read.
Will the titles line up automatically? They can. Vizard Agent places them on the segment boundaries it recorded.
How is the music handled? Vizard Agent measures it against the narration rather than setting it by ear, so the words stay on top.
Can I hear the segments before the full build? Yes. Ask Vizard Agent for the narration track on its own first.
What if I want a different voice for one section? Vizard Agent can, though it usually reads as a second narrator rather than a change of mood.
Can it write the direction for me? Yes, from the script. You should still check where Vizard Agent decided the turns were.
Does a clone work the same way? Yes, with the same segmenting, though a clone has less range than a designed voice.
What if the script changes later? Only the affected segments need regenerating, so Vizard Agent leaves the rest of the read alone.