How to change someone's face in a video without the voice drifting out of sync
Ask Vizard Agent to test the change on a single shot with two different models before touching the rest, and to keep the original audio separate so it can be re-aligned afterwards. Changing a face is not a grade or an overlay — the frames are remade, and remade frames do not necessarily arrive on the same timing as the ones they replaced.
What is the short version?
Softening a face, removing a beard or changing someone's apparent age all mean the same thing underneath: the picture gets generated again from scratch. The voice was not regenerated along with it, so the two tracks can end up describing slightly different moments.
- Ask Vizard Agent to test one shot before doing all of them.
- Ask it to keep the original audio aside for re-alignment.
- Ask it to check the gestures still match the words afterwards.
What do you need before you start?
The video and a decision about how far you want to go. Vizard Agent can do a light retouch that leaves the timing alone or a real change that requires regeneration, and the difference between them is exactly the difference between keeping sync and having to rebuild it.
- The video. With its original audio.
- What should change. Beard, age, blemish, expression.
- How strong. Light retouch, or an actual change.
- Which shots. All of them, or some.
- Whether sync is critical. It usually is for speech.
What do you type into Vizard Agent?
Ask Vizard Agent for the test before the job. A change judged on a single shot at full size costs a small fraction of doing the whole video, and it tells you whether the approach is viable at all before anybody spends time repairing sync.
Prompt
Variants worth knowing:
- "Try it on one shot first." The cheap decision.
- "Two different approaches." They fail differently.
- "Tell me what that does to the timing." Names the real risk.
What does Vizard Agent actually do?
Here is the sequence on a real piece where the subject's face had to be changed across several shots. The order matters: the comparison happens before the commitment, and the sync check happens after every pass rather than at the end.
- Reviews the video and extracts frames of the subject.
- Reads the transcript so the speech timing is known.
- Tests the facial change with two different models on one shot.
- Compares the two results side by side.
- Zooms into both faces to judge them at real detail.
- Tests consistency between blocks of the regenerated footage.
- Verifies the result is still in sync with the audio.
- Checks the gestures still match what is being said.
- Applies the chosen treatment to every shot of the subject.
- Compares consistency across shots, not just within one.
Step three is worth the extra pass every time. Two models given the same face produce different failures — one softens the jaw, another changes the eyes — and the only way to know which you can live with is to look at both.
Step ten is the trap in multi-shot work. Each shot is regenerated independently, so the face can be slightly different in each one, and a comparison within a single shot will never show you that.
Step eight is not the same as step seven. Sync is whether the mouth matches the words; gestures are whether the hand still lands on the emphasis — and a regeneration can preserve one and quietly shift the other.
What does the result look like?
The change you asked for, with the voice still attached to it. Vizard Agent tells you what the regeneration did to the timing and what it did to re-align the audio, so the sync is a reported fact rather than something you check by squinting.
When does this not work well?
Some changes cost more than they buy. If the subject is talking in close-up throughout, every frame of every shot has to be regenerated, and the accumulated drift and per-shot inconsistency can be worse than the thing you were fixing.
- Close-up speech throughout. Every frame is at risk.
- Long takes. Drift accumulates across the shot.
- Many shots of one person. Each regenerates differently.
- Subtle changes. Often a light retouch is enough and keeps timing.
- Critical lip sync. A dub or a cutaway may be safer.
How do you fix a result that came back wrong?
Separate the three failures before you report them. "It does not look like him", "the mouth is off the words" and "he looks different in shot four" have three entirely different causes, and Vizard Agent treats each one as its own fix rather than re-running everything.
- "It does not look like him." The other model is tried on that shot.
- "The mouth is off the words." The audio is re-aligned to the new timing.
- "He changes between shots." Consistency is compared across all of them.
- "The gesture lands late now." The gesture timing is checked separately from sync.
How does Vizard Agent compare to doing it yourself?
By hand you run the face change across the whole video because that is how the tool works, and then discover the sync is off. Fixing it means either re-aligning audio that no longer matches anything exactly, or running the whole thing again with different settings.
| By hand | Vizard Agent | |
|---|---|---|
| Testing | On the finished video | On one shot, two ways |
| Judging it | At playback size | Zoomed in on the face |
| The audio | Baked in and broken | Kept aside, re-aligned |
| Between shots | Assumed consistent | Compared across all of them |
| Gestures | Not checked | Checked separately from sync |
Common questions
Why does the sync break at all? Because the picture is regenerated, not adjusted. Vizard Agent keeps the audio aside for that reason.
Is a light retouch safer? Much. Vizard Agent will tell you when the lighter route gets you what you want.
Can it try two approaches? Yes, and it is worth it. Vizard Agent shows you how each one fails.
Will he look the same in every shot? Only if it is checked. Vizard Agent compares across shots as well as within one.
What about the lip sync specifically? Verified after every pass. Ask Vizard Agent for the check, not just the file.
Can it fix the sync afterwards? Usually, by re-aligning the audio. Vizard Agent reports what it had to shift.
What is a gesture check? Whether the hands still land on the emphasis. It can drift while the mouth does not.
Does this apply to age changes? Yes, and more so. A bigger change means more regeneration.
Can it do just one shot? Yes. Say which, and Vizard Agent leaves the rest untouched.
How long are the safe takes? Shorter is safer. Vizard Agent flags where drift starts accumulating.
Will the background change too? It can. Say what must stay and Vizard Agent checks it.
Can I keep the original as well? Yes, and you should. Vizard Agent keeps the comparison available.
Is a cutaway ever the better answer? Often, for a single problem shot. Vizard Agent will say so.
What if I only care about stills? Then none of this applies — the timing problem is unique to moving pictures.