Vizard Agent

How to slow down a narrator without breaking the lip sync

Last updated 2026-09-04 · 7 min read

Tell Vizard Agent to slow the delivery and to keep it looking natural. It slows the voice and, for anyone on camera, slows the picture with them rather than only the audio — then re-times the on-screen graphics to the new pace and checks the mouths frame by frame before delivering.

What is the short version?

"Can you slow her down a little" sounds like a volume knob and is not one. Slowing a voice moves every word later, which pulls the mouth out of step, drags the graphics out of position, and changes the length of the whole cut.

  1. Go to Vizard Agent with the current version.
  2. Say who is talking too fast and roughly how much slower.
  3. Say the lip sync must not break and it has to look natural.

What do you need before you start?

The delivered cut and a sense of how much slower you want it. Vizard Agent handles the consequences — the sync, the graphics, the runtime — but "a little slower" is genuinely enough of a brief to start from, and easier to react to than a percentage.

What do you type into Vizard Agent?

Say the constraint out loud rather than assuming it. Without it, the reasonable reading of "slow her down" is to stretch the audio track — which works perfectly well for a voiceover and looks immediately wrong on anyone whose face happens to be on screen while they speak.

Prompt

The client says the narrator talks too fast — can you slow her down a little? Everyone should be slower. But I don't want the lip sync to break, it needs to look natural.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real insurance ad that came back too fast for the client to approve. Step three is the one that separates a natural-sounding result from a rubbery one, and it is the step most people skip entirely.

  1. Slows the narration and re-times the whole cut around it.
  2. Checks the new mix and the pace of the result.
  3. Slows the on-camera presenters together with their picture.
  4. Fixes the frame-exact lengths and rebuilds the timeline.
  5. Inspects the presenters' faces for stretching artifacts.
  6. Verifies the lip sync on the slowed presenters specifically.
  7. Measures the graphic and word timings against each other.
  8. Checks when each on-screen list item appears.
  9. Re-times the on-screen list to the new spoken line.
  10. Verifies the list matches the words on real frames.
  11. Smooths a shot with optical flow where the slow-down showed.
  12. Checks the interpolated frames for artifacts.
  13. Re-measures the levels and confirms the closing mix.

Step three is the whole method. A voiceover can be slowed on its own because there is no mouth to contradict it. A presenter cannot — the picture has to be slowed by exactly the same factor, which is why "slow the narrator" and "slow the presenters" are two different operations inside the same request.

Steps seven to ten are the part people forget until they watch it back. Every bullet, badge and list item was timed to a word that has now moved, and a checklist that appears half a second before the line that introduces it reads as a mistake even when nobody can say what is wrong.

What does the result look like?

A cut at a comfortable pace where nobody sounds stretched, the presenters' mouths still match what they say, and the graphics still land on the lines they belong to. It runs longer than the original, because slowing something down is a thing that takes time.

Where the slow-down would have shown as judder, the picture is interpolated rather than left to stutter — and those frames are checked, because interpolation has its own tells.

When does this not work well?

There is a limit to how far speech can be slowed before it stops sounding like a person talking. Past roughly a tenth, the ear starts to notice, and past a fifth it is obvious no matter how good the processing is.

How do you fix a result that came back wrong?

React to what you hear or see. Vizard Agent keeps the timing map, the graphic positions and the sync verification, so a further adjustment re-times from the same analysis rather than stretching an already-stretched track a second time.

How does Vizard Agent compare to doing it yourself?

By hand the audio part takes a minute and everything downstream takes the afternoon: retime the picture, retime the graphics, fix the music, then discover that one presenter is now visibly out and start again on that shot.

By hand Vizard Agent
Slowing the voice One operation One operation
On-camera speakers Often forgotten Picture slowed with them
Graphics Manually nudged Re-timed to the new words
Checking the sync Play it back Verified on frames
Judder Live with it Interpolated, then inspected

Common questions

How much slower can it go? Around a tenth before it becomes audible. Vizard Agent will say when it is pushing it.

Does the video get longer? Yes. Slower speech takes more time; there is no way around that.

What if the runtime is fixed? Then something has to be cut. Vizard Agent will tell you what it recommends.

Can it slow only one speaker? Yes. Name them and Vizard Agent leaves the others alone.

Will the presenter look stretched? No. Vizard Agent slows the picture with the voice rather than only the audio.

What about the music? Re-fitted. Vizard Agent adjusts the bed to the new length.

Does it re-record instead? Vizard Agent can, for a generated voice. Say which you prefer.

Will the captions still match? Yes. Vizard Agent re-derives them from the slowed cut.

Can it speed things up too? Yes, and the same rules apply in reverse.

What is optical flow for? Smoothing a slowed shot so it does not judder. Vizard Agent checks the result for artifacts.

Can I ask for a specific percentage? Yes, though telling Vizard Agent "a little slower" usually gets there in fewer rounds.

Does the mix change? Vizard Agent re-measures the levels afterwards, since the timing moved.

What if only one shot looks wrong? Say which one and Vizard Agent re-treats only that shot.

Why not just slow the audio? Because anyone on camera will immediately look dubbed, and that is a bigger problem than the pace was.