How to change a few spoken words in a video
Upload the video to Vizard Agent and say which words should change and what they should become. Vizard Agent finds the phrase in the word-level transcript, clones the speaker's voice for the new wording, lip-syncs just those seconds, and cuts on a silent point in the waveform so the repair does not announce itself.
What is the short version?
One wrong date in an otherwise finished video normally means filming it again. The alternative is replacing a second and a half of speech convincingly, which means the voice has to match, the mouth has to match, and the join has to be inaudible.
- Go to Vizard Agent and upload the video.
- Say exactly which words are wrong and exactly what they should say instead.
- Say whether any text on screen shows the same information.
What do you need before you start?
The video and the precise wording. Vizard Agent locates the phrase itself from the transcript, so what it needs from you is the replacement text written out exactly as it should be spoken, including the form of a date or a number, because "the seventeenth" and "August seventeenth" are different lengths of speech.
- The video with clean audio. The clearer the original speech, the better the clone matches it.
- The old phrase and the new one. Written out exactly. Ambiguity here is expensive.
- Whether the speaker is on camera. If their mouth is visible the segment needs lip-syncing as well.
- Any on-screen text saying the same thing. A burned-in date has to change with the audio.
- Whether you have permission. Changing what someone said on camera is a decision about them, not about the file.
What do you type into Vizard Agent?
Quote both phrases. Vizard Agent finds the moment from the transcript rather than needing a timestamp from you, and giving it the exact replacement wording is what lets it generate speech that fits the gap it is going into.
Prompt
Variants worth knowing:
- Audio only. If the speaker is off camera or looking away, no lip-sync is needed and the job is much simpler.
- Several corrections at once. Cheaper than one at a time, since the transcript and the voice sample are done once.
- Remove rather than replace. Cutting a phrase out cleanly is a related and often easier request.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real word replacement. Notice how much of it is verification: it generated the new speech twice, measured the audio energy to choose the cut points, and hunted down the on-screen text frame by frame.
- Transcribes the video to word-level timings and reads the file to find the exact phrase.
- Analyses the speech for pauses and stutters around the target.
- Extracts the audio as a clean sample and uploads it as the voice reference.
- Clones the speaker's voice for the new phrase with two different engines and compares them.
- Probes and analyses the generated clips for length and naturalness.
- Finds the coordinates of the on-screen text box and crops it to confirm the contents.
- Scans the video at 0.1-second intervals and tiles the crops to find exactly when that text appears and disappears.
- Lip-syncs only the affected seconds to the new audio, then analyses the result.
- Measures audio energy around the join to pick the quietest point to cut on.
Step 9 is the difference between a repair and an obvious edit. Vizard Agent looked for the sub-millisecond low point in the waveform rather than cutting on a round number, which is what stops the join from clicking.
What does the result look like?
From the run this page is written from, probed on the delivered file: 744x1280, H.264, 60fps, 17.26 seconds, AAC audio. The same length as the original, at the same frame rate, with a second and a half of new speech inside it and the on-screen text rebuilt in a matching font.
Vizard Agent keeps the untouched audio exactly as recorded — only the replaced span is generated, so the rest of the video is still your original voice rather than a re-read.
Timing was not measured separately for this kind of job. The closest measured work runs a median of 28 to 38 minutes end to end, with the middle half spread considerably wider. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.
When does this not work well?
The technical limits here are real, and they are the smaller half. Replacing someone's recorded words changes the record of what they said, and Vizard Agent will do exactly what you ask without knowing whether the person in the video agreed to it.
- It has to be your own video, or you need permission. This is the limit that matters most.
- A much longer replacement will not fit. New speech that runs well past the old span forces the timing around it to shift.
- Noisy or reverberant originals clone poorly. The cleaner the source voice, the closer the match.
- A visible mouth raises the bar. Lip-sync is good and a close-up of a face is where any imperfection shows first.
- Regulated claims are still claims. Changing a price or a date in an advertisement does not make the new one compliant. That check is yours.
How do you fix a result that came back wrong?
Say what sounds or looks off. Vizard Agent keeps the cloned takes, the word timings, the energy measurements and the text crops, so trying a different voice engine or shifting the join by a few milliseconds does not repeat any of the analysis.
- "The new words sound too fast." Vizard Agent regenerates at the original speaking pace.
- "I can hear the join." A different cut point from the same energy measurements.
- "The on-screen text is the wrong font." Rebuilt against the fonts it matched from the frames.
How does Vizard Agent compare to doing it yourself?
By hand this is a re-record that never quite matches the room, or a splice you can hear. The professional route is to bring the speaker back and film the line again, which costs a day and is usually why the video simply ships with the mistake in it.
| By hand | Vizard Agent | |
|---|---|---|
| The new speech | Re-record, mismatched room | Cloned from the original audio |
| The mouth | Cut away and hope | Lip-synced on those seconds only |
| The join | Cut on a round number | Cut on the measured low point |
| On-screen text | Find it by scrubbing | Located frame by frame |
Common questions
Will it sound like me? It clones from your own audio in the video, so it matches closely when the original recording is clean.
Does the whole video get regenerated? No. Vizard Agent replaces only the affected span and leaves the rest of your recording untouched.
Can it change text shown on screen too? Yes. Vizard Agent locates the text box, works out exactly when it is visible, and rebuilds it in a matching font.
What if the person is not facing the camera? Then no lip-sync is needed and Vizard Agent replaces the audio alone, which is both simpler and faster.
Can I change several phrases in one video? Yes, and doing them together is cheaper than one request each, because Vizard Agent transcribes and samples the voice once for all of them.
Is this allowed? For your own footage, yes. Putting new words in someone else's mouth is a decision that needs their agreement, and Vizard Agent cannot make it for you.