How to lip-sync generated animation to a recording without moving the audio
Tell Vizard Agent the original recording must not be cut or moved. It plans the animation scene by scene against the transcript, transcribes each generated clip to match its words to the recording's, and stretches or compresses only the picture between those anchor words so every mouth lands on the original audio.
What do you type into Vizard Agent?
Make the audio the fixed point and ask for a short preview. A recorded conversation's comedy and meaning live in its timing, and the cheap way to fix drift is to shift the sound to the picture, which is exactly what you have ruled out. A preview of a few shots is the cheapest place to find out whether the look works.
Turn this recorded call into a 3D cartoon. Use the exact original audio and keep all the dialogue, laughter and timing. Sync the characters' lips and reactions to the real conversation. Show me a short preview first.
What does Vizard Agent actually do?
It treats each generated clip as approximate and the recording as truth. A video model given the lines and voice samples gets the words roughly right but never to the frame, so the alignment happens afterwards, word by word. In the session behind this article it:
- Transcribed the four-minute recording with word timings and planned 21 scenes, each prompt carrying its exact lines and timings plus voice samples of each speaker.
- Delivered a three-shot, 16-second preview for approval before generating the rest.
- Transcribed every generated clip and matched its words to the original, reporting a match rate per scene and fixing a low score caused by a stammered word.
- Turned the matched words into anchor points and re-timed only the picture between them to exact frame counts, confirming the preview's audio was effectively identical to the source.
- Reviewed sync phrase by phrase at ten frames a second and replaced shots with stray objects or duplicated characters.
When does this not work well?
When the speakers talk over each other. Overlapping speech gives the transcript two sets of words in the same moment, and the anchors become ambiguous. Very fast exchanges also force large speed changes on short shots, which can make the animation look jerky between anchors.
Common questions
Why does my animated version drift out of sync? Generated clips are approximate. Vizard Agent re-times the picture to the words.
Will the original audio be changed? No. Vizard Agent keeps it exactly as recorded.
Can I approve the look first? Yes. Vizard Agent makes a short preview before the full video.