How to stop an avatar making faces the real person never makes
Tell Vizard Agent the face is right but the expressions are not. It goes back to your source video, studies the range the person actually uses, maps where their face is unobstructed second by second, builds a base cut from that clean footage and lip-syncs it to the new script — so the expressions are theirs rather than invented.
What is the short version?
A generated presenter usually gets the likeness right and the behaviour wrong. It nods more, smiles wider and raises its eyebrows in places the real person never would, and everyone who knows them notices immediately. Their own footage does not have that problem.
- Go to Vizard Agent with the avatar and the source video.
- Say the expressions are not how that person behaves.
- Ask for their real footage lip-synced instead.
What do you need before you start?
The original video and enough of it. What matters is not its length but how much of it shows the face clearly — unobstructed, facing forward, without a hand or a microphone across it — because that is the material the new version is built from.
- The source video. As long as you have.
- What is wrong. Expressions, not likeness.
- The new script. Or the voice track for it.
- How they actually behave. Say it if you can.
- The format you need. It affects the framing.
What do you type into Vizard Agent?
Describe the person rather than the fault. "He never pulls those faces — go and watch the original" points at evidence, and evidence is something that can be studied; "make it more natural" is a direction with no target.
Prompt
Variants worth knowing:
- "Rewatch the original." Names the reference.
- "His own footage, lip-synced." The method that fixes it.
- "He never does that." A fact, not an opinion.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real spokesperson video whose generated avatar was performing considerably more than its subject ever did. The turn in this list happens at step four, where the whole approach changes from generating a face to using the one already on file.
- Maps where the face is unobstructed in the original video.
- Studies the real expression style across that footage.
- Checks face visibility second by second to find usable spans.
- Assembles a base cut from the clean, unobstructed footage.
- Reviews the assembled base cut before syncing anything.
- Lip-syncs the real footage to the new voice track.
- Measures where the head sits in each segment.
- Reframes the base cut tighter on the face.
- Lip-syncs the reframed footage and inspects the result.
- Zooms into the mouth region to check the sync closely.
- Finishes with captions and music over the new cut.
- Has the finished version watched back end to end.
Step three is the unglamorous step that makes it possible. A video that seems to show somebody talking for two minutes usually contains far less usable face than that — a hand goes up, they look away, something crosses the frame — and the base cut can only be built from what survives that check.
Step seven is why the result holds together. Cutting between spans from different parts of a recording puts the head in a slightly different place each time, so each segment is measured and the framing is re-derived rather than one crop being applied to all of them.
What does the result look like?
The same person, saying the new script, behaving the way they actually behave on camera: their own eye movements, their own stillness, their own small habits, with the mouth Vizard Agent synced matched precisely to the new words on the new track.
The uncanny quality goes away entirely, because there is nothing left in the picture for it to attach to.
When does this not work well?
You need clean footage of the face. If the source is short, badly lit, mostly in profile, or has something across the face for most of its length, there may not be enough to build a base cut from — and then a generated avatar really is the only option.
- Very little clean footage. Nothing to assemble.
- Constant movement. The lip-sync has less to hold on to.
- Profile shots. The mouth has to be visible.
- A long new script. Short source footage has to repeat.
- A different look required. Their own footage is what it is.
How do you fix a result that came back wrong?
Say what you see rather than what you would change. Vizard Agent keeps the visibility map and the base cut, so a note about one segment swaps that span for another piece of clean footage instead of rebuilding the video.
- "That segment looks odd." Replaced from another clean span.
- "The mouth is off." The sync re-checked at the mouth.
- "Too tight a crop." Reframed from the head measurements.
- "It cuts too often." Longer spans preferred where they exist.
How does Vizard Agent compare to doing it yourself?
The usual response to a stiff or overacting avatar is to regenerate it and hope for a better take, which changes the performance randomly rather than towards the person. Going back to their footage is more work and is the only approach that converges.
| By hand | Vizard Agent | |
|---|---|---|
| The fix | Regenerate and hope | Use the person's own footage |
| The reference | An impression | The source video, studied |
| Usable material | Assumed | Mapped second by second |
| Framing across cuts | One crop | Measured per segment |
| Checking the sync | Watching | Inspected at the mouth |
Common questions
Does it still work as an avatar? It is better than one. Vizard Agent uses the real person.
How much source footage is needed? More than you expect, because Vizard Agent counts only clean face time.
Can it use several videos? Yes. Give Vizard Agent everything you have of them.
Will the voice be theirs? It can be, if Vizard Agent has clean audio to clone from.
What if they look away a lot? Vizard Agent works around it and tells you what it found.
Can it keep their gestures? Yes. That is the point of Vizard Agent using their own footage.
What about a longer script? Vizard Agent will need more source or will reuse spans.
Does the lip-sync look convincing? Vizard Agent checks it zoomed in on the mouth.
Can it change the background? Yes, though Vizard Agent finds that reintroduces some artificial feel.
Will it match my brand look? Yes. Vizard Agent applies captions, colour and music over it.
Can I approve the base cut first? Yes, and it is worth doing before any syncing.
Does this work for someone who is camera-shy? Yes, and Vizard Agent usually needs less footage than they fear.
What if I only have photographs? Then a generated performance is the only option Vizard Agent has.
Why not just prompt for calmer expressions? Because a description of somebody's manner is a guess, and their own footage is a record of it.