Vizard Agent

How to replace what someone says with a new script in their own voice

Last updated 2026-08-30 · 7 min read

Upload the video and paste the script you want said instead. Vizard Agent analyses your delivery, energy and room acoustics, clones the voice, tests more than one engine before choosing, and lines the new speech up so its emphasis lands on the same physical moments as the original.

What is the short version?

Re-recording a piece to camera means the lighting, the setting and the performance all change with it. Having Vizard Agent replace only the words keeps everything that was right about the original take and fixes the single thing that was not.

  1. Go to Vizard Agent and upload the video of yourself.
  2. Paste the new script exactly as it should be said.
  3. Say it must be your own voice and your own delivery.

What do you need before you start?

Your video and the new wording. Vizard Agent takes the voice model from the recording itself, so no separate sample is needed from you, and it studies the delivery closely so the replacement carries the same energy rather than reading flat.

What do you type into Vizard Agent?

Paste the script and say it is your own voice. Vizard Agent treats a voice-replacement request on the speaker's own footage differently from a generic narration job, and the delivery analysis it does only happens when the performance is meant to be preserved.

Prompt

Please change what I say in this video to: "[the new script, in full]". Keep my own voice and the same delivery and energy as the original.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real full-script replacement. The step that makes it convincing rather than merely accurate is the twentieth: it found the exact frame where the speaker slammed the table, so the new emphasis lands there too.

  1. Inspects the available tools and probes the input video's audio and lip-sync options.
  2. Downloads the video, transcribes the spoken words and extracts the audio.
  3. Samples the video frames and reviews the speaker, the setting and the visuals.
  4. Reads the detailed word timings from the transcription.
  5. Locates and prints the full word timestamps for the original delivery.
  6. Analyses the speaker's delivery, energy, room acoustics and sound effects.
  7. Tests voice cloning on two separate engines and analyses the generated samples.
  8. Generates the full speech on both engines and analyses each for dynamic range, emotion and escalation.
  9. Analyses the audio energy around the table slam, extracts frames at that moment and pinpoints the exact instant of the physical gesture.

Step 6 is what most people would skip. Room acoustics are part of how a voice sounds in a given space, and a cloned voice generated without them lands on top of the footage rather than inside it.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1280x720, H.264, 30fps, 21.73 seconds, AAC audio. Twenty-two seconds of the same person in the same room with the same gestures, saying something different, in a voice that is recognisably theirs and builds to the same physical moment.

The picture is untouched. That is the point of the method: everything the original take got right stays exactly as it was.

When does this not work well?

Putting new words into a real person's mouth is a powerful capability and precisely the one that must never be misused. These limits are not optional, and Vizard Agent states them plainly rather than leaving them implied in the small print.

How do you fix a result that came back wrong?

Say what sounds or looks off to you and where. Vizard Agent keeps both engines' samples, the word timings, the full delivery analysis and the gesture frames, so the voice or the alignment can change without starting again.

How does Vizard Agent compare to doing it yourself?

By hand this means setting up and filming the whole thing again, in different light, on a different day, with a performance you will probably like less. The alternative is a voiceover that visibly does not match the mouth.

By hand Vizard Agent
The picture Re-shot, and different Kept exactly as it was
The voice Re-recorded, or replaced badly Cloned and compared across engines
Delivery Perform it again Energy and escalation analysed and matched
Gestures Hope they still fit Emphasis aligned to the actual frames

Common questions

Can I do this to someone else's video? Only with their explicit consent. This is not a technique to use on a person who has not agreed.

Will it sound like me? Closely. Vizard Agent tests more than one cloning engine and chooses the better match.

Does the picture change? No. Vizard Agent replaces the audio and aligns it; the footage is untouched.

What if the new script is much longer? Then the mouth will not match. Keep the length close to the original.

Does it match my gestures? It does. Vizard Agent finds the physical moments and lands the new emphasis on them.

Why analyse the room acoustics? Because a voice recorded in a room carries that room, and a clean cloned line dropped into the same footage sounds like it was added afterwards — which it was, and which is exactly what you do not want anyone to notice.

Should I tell people? Where it matters, yes. Some contexts and jurisdictions require disclosure.

Can it keep the original energy? Yes. Vizard Agent measures the dynamic range and the escalation and reproduces both.

How long can the script be? As long as the original delivery. Beyond that the lip movements stop working.

When is re-shooting the better answer? Whenever the new script is much longer or says something quite different. Vizard Agent matches a script of similar length and energy well, and beyond that the mouth stops agreeing with the words.

Can I fix just one sentence instead? Yes, and it is the safer option. Vizard Agent replaces a single line as readily as a whole script, and a smaller change is far less likely to fight the lip movements.

Why test two cloning engines rather than one? Because voices clone unevenly. Vizard Agent generates the full speech on both and compares them for dynamic range and emotion, since the engine that wins on one voice regularly loses on the next.

Does Vizard Agent check the result? Yes. It reviews the alignment against the gesture frames before delivery.