Vizard Agent

How to match how someone speaks and not just how their voice sounds

Last updated 2026-09-15 · 8 min read

Tell Vizard Agent to match the delivery as well as the voice, and let it choose its reference from the liveliest part of the recording rather than the cleanest. It compares the new read against the original at matched volume, checks pitch variation and emphasis rather than timbre alone, then fits the edit around where the new read breathes.

What is the short version?

A voice match has two halves. The first is timbre, which is what people mean by cloning and what comes out well from any clean sample. The second is manner — pace, emphasis, where the pauses fall — and a clean sample teaches none of it.

  1. Point Vizard Agent at an expressive part of the recording.
  2. Ask for the delivery matched, not just the voice.
  3. Let the edit move to fit the new read.

What do you need before you start?

A recording of the person and a description of how they speak when they are being themselves. Vizard Agent can hear the voice, but whether a read is characteristic of someone is a judgement only people who know them can make.

What do you type into Vizard Agent?

Describe the delivery in ordinary words. A phrase like "conversational, varied pitch, emphasis on the key words, pauses that mean something" gives Vizard Agent something specific to aim at and, more usefully still, something concrete to check the finished read against afterwards.

Prompt

Here is a recording of him. Match his voice, but also the way he talks — expressive and conversational, natural pacing, varied pitch, real emphasis on key words and pauses that are deliberate rather than gaps. Use the most animated part of the recording as your reference, not the calm opening, and check the firm's name against how he says it.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on a real advert re-voiced to match one particular person. It takes three passes in all, and the interesting one is the second, where a technically accurate voice match was rejected for sounding like somebody reading aloud.

  1. Listens to the reference and checks it is one clear voice.
  2. Builds the voice match and times it to the existing scenes.
  3. Checks how closely the read matches in tone and pacing.
  4. Refines it after the first attempt comes back rushed and synthetic.
  5. Compares the refined read against the reference at matched volume.
  6. Checks how the person says a difficult name in the original recording.
  7. Applies a phonetic guide for that name without changing the script.
  8. Finds the most expressive passages of the recording to steer the delivery.
  9. Re-reads using that animated passage as the reference.
  10. Fits the edit around the new pauses rather than stretching the read.

Step five sounds pedantic and is not. A louder read sounds more confident, so comparing two takes at different volumes tells you which is louder rather than which is closer.

Step eight is the finding worth remembering. The sample you would choose for a clean clone — calm, evenly paced, no laughter — is the sample that produces a flat read, because that is what it was taught.

Step ten decides whether it sounds real. A read squeezed into the old timings loses exactly the pauses that made it sound like a person, so the edit moves and the read stays as delivered.

What does the result look like?

A read that sounds like the person rather than merely like the person's voice. The emphasis falls where they would put it, the pauses mean something, and Vizard Agent hands over the comparison against the reference along with a note of where it moved the edit to accommodate the new pauses.

When does this not work well?

Manner is much harder to borrow than tone. A reference that is short, flat, or recorded with music underneath it gives Vizard Agent very little to learn from, and a script written in nobody's particular voice will sound stilted no matter who ends up reading it.

How do you fix a result that came back wrong?

Say what is wrong with the performance rather than with the voice. "It sounds rushed" and "it does not sound like him" send Vizard Agent to two completely different places — the first to the pacing of the read, the second back to the choice of reference sample.

How does Vizard Agent compare to doing it yourself?

By hand you pick the cleanest thirty seconds as the sample, because that is what the guidance says, and get back a perfect impression of somebody reading aloud. The fix is counter-intuitive enough that most people try a different tool instead.

By hand Vizard Agent
The sample The cleanest passage The most expressive one
What is matched Timbre Timbre and delivery
Comparing takes By ear, at whatever volume At matched volume
Difficult names Guessed or spelled out Taken from the recording
Timing Read stretched to the edit Edit fitted to the read

Common questions

Why does a clean sample sound flat? Because it is flat. Vizard Agent looks for an animated passage instead.

How long does the reference need to be? Long enough to contain some range. Vizard Agent will say if it is too short.

Can it match someone laughing or emphatic? It can lean that way, and Vizard Agent picks the sample accordingly.

What about a name it gets wrong? Vizard Agent checks the original recording for how the person says it.

Will the timing still fit the picture? Vizard Agent adjusts the edit around the new pauses rather than compressing them.

Can I hear the comparison? Yes. Vizard Agent compares the new read with the reference at matched volume.

What if there is music under the reference? Separate it first. Vizard Agent can do that before building the match.

Can it do several lines in different moods? Yes. Tell Vizard Agent which lines and it steers each one.

Does the script change? No, unless you ask. Only the delivery and the pronunciation guide change.

What if I only have a phone recording? Usually enough. Vizard Agent cares more about range than about fidelity.

Will the captions still sync? Yes. Vizard Agent resynchronises them to the new read.

Can it keep some of the original audio? Yes. Say which parts and Vizard Agent leaves those alone.

How many attempts does it take? Often two or three, because the first match is judged and then refined.

Is a flat read ever right? Sometimes — for a formal read. Then the clean sample is the correct choice.