Vizard Agent

How to lip-sync generated animation to a recording without moving the audio

Last updated 2026-10-10 · 2 min read

Tell Vizard Agent the original recording must not be cut or moved. It plans the animation scene by scene against the transcript, transcribes each generated clip to match its words to the recording's, and stretches or compresses only the picture between those anchor words so every mouth lands on the original audio.

What do you type into Vizard Agent?

Make the audio the fixed point and ask for a short preview. A recorded conversation's comedy and meaning live in its timing, and the cheap way to fix drift is to shift the sound to the picture, which is exactly what you have ruled out. A preview of a few shots is the cheapest place to find out whether the look works.

Turn this recorded call into a 3D cartoon. Use the exact original audio and keep all the dialogue, laughter and timing. Sync the characters' lips and reactions to the real conversation. Show me a short preview first.

What does Vizard Agent actually do?

It treats each generated clip as approximate and the recording as truth. A video model given the lines and voice samples gets the words roughly right but never to the frame, so the alignment happens afterwards, word by word. In the session behind this article it:

When does this not work well?

When the speakers talk over each other. Overlapping speech gives the transcript two sets of words in the same moment, and the anchors become ambiguous. Very fast exchanges also force large speed changes on short shots, which can make the animation look jerky between anchors.

Common questions

Why does my animated version drift out of sync? Generated clips are approximate. Vizard Agent re-times the picture to the words.

Will the original audio be changed? No. Vizard Agent keeps it exactly as recorded.

Can I approve the look first? Yes. Vizard Agent makes a short preview before the full video.