Vizard Agent

How to stop a screen recording stalling under the voiceover

Last updated 2026-09-04 · 8 min read

Send the screen recording and the separately recorded voiceover, and give Vizard Agent two rules: show whatever the narration is talking about, and never let the picture sit still for more than a few seconds. It finds the stalls and fills them with keyword cards, magnified detail or related shots.

What is the short version?

A tutorial built from narration plus a screen capture has one very specific failure mode: the voice keeps explaining while the screen sits on a static settings panel. Nothing about it is technically wrong, every word is correct, and the viewer leaves anyway.

  1. Go to Vizard Agent with the recording and the voiceover.
  2. Say the picture must show what the narration is describing.
  3. Say how long the screen may sit still before something must happen.

What do you need before you start?

Both files, and a number. The number is the important part — "don't let it get boring" is not actionable, and "never more than five to eight seconds without visible movement" is a rule that can be applied and checked across a whole video.

What do you type into Vizard Agent?

Give the rule and the exception together. The instruction that carries the whole job is not "make it engaging" — it is that when the screen stops moving, something else has to, and you say how long the picture is allowed to wait first.

Prompt

Cut this screen recording to my voiceover. Keep them exactly in step — show whatever I'm talking about at that moment. If the screen sits still for more than 5–8 seconds, cut in something related: B-roll, a keyword animation, or a zoom on the part I'm describing. Keep each insert short.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real product tutorial built from a screen capture and a separate narration. Steps two and three are what makes the rest possible: it has to know exactly when each action happens.

  1. Checks both files and loads the tutorial editing approach.
  2. Analyses the narration words against the screen picture in parallel.
  3. Samples frames densely to locate the precise moment of each action.
  4. Detects the window and editor boundaries in the capture.
  5. Measures the crop boundaries for each region of the interface.
  6. Renders the segments against the narration's line, not the recording's order.
  7. Checks each segment's opening frame for composition.
  8. Renders a progress bar and section titles over the top.
  9. Builds the opening title and keyword animations, and lays in the voiceover.
  10. Splits a section into a wide shot and a close-up where one shot was not enough.
  11. Analyses the assembled cut for sync and flow.
  12. Adds an end card reviewing the steps.
  13. Diagnoses a mis-crop and a missing element found on review.
  14. Builds a magnifier effect aligned to one specific line on screen.
  15. Measures the preview area's exact bounding box so the magnifier lands.
  16. Re-renders the magnifier until it sits on the right line.
  17. Adds keyword capsules timed to the narration.
  18. Adds a slow push-in to the section that was otherwise static.
  19. Diagnoses a repeated second and a jump cut and rebuilds those.
  20. Verifies the progress bar advances and the ending holds.

Step six is the reordering that people do not expect. The screen recording happened in the order you clicked; the narration explains things in the order that makes sense — so the picture is cut to the voiceover's sequence rather than played back as recorded.

Steps fourteen to sixteen are what a stall gets filled with when there is nothing else to show. The screen was static on one panel, so a magnifier was built over the exact line being discussed — measured against the real bounding box, because a magnifier over the wrong row is worse than no magnifier.

What does the result look like?

A tutorial where the picture is always doing something, and always doing the specific thing being described at that moment. Where the interface genuinely had nothing left to show, Vizard Agent carries those seconds with a keyword card, a slow push-in or a magnified detail instead.

The inserts are short by design. A stall filled with a ten-second stock clip is a different problem wearing the same clothes — the fix is punctuation, not a second video running underneath the first.

When does this not work well?

The method needs the recording to actually contain the thing the narration is describing. Where the voiceover talks at length about a screen that was never captured, no amount of pacing work by Vizard Agent invents that screen, and the honest answer is to record it.

How do you fix a result that came back wrong?

Point at the second where it goes wrong. Vizard Agent keeps the measured action timings, the interface region boundaries and the narration alignment from the first pass, so a mis-timed insert or a wrong crop becomes a rebuild of that one segment rather than of the whole tutorial.

How does Vizard Agent compare to doing it yourself?

By hand this is the most tedious kind of editing there is: scrub to find where you clicked, match it to where you said it, and then decide what to do about the forty seconds where you explain something while the screen shows a settings panel.

By hand Vizard Agent
Finding each action Scrubbing Located by dense sampling
Matching to the narration By ear, clip by clip Cut to the voiceover's order
Stalls Noticed on playback Found against a stated limit
Filling them Whatever is to hand Keyword cards, zooms, magnifiers
Checking Watch it through Sync and flow analysed

Common questions

Do I record the voiceover first? Either way. Vizard Agent cuts the picture to the voice, so voice-first is easier.

What is a good stall limit? Five to eight seconds. Vizard Agent applies whatever number you give it.

Will it reorder my recording? Yes, to follow the narration. Say so if the order must be preserved.

Can it zoom instead of cutting away? Yes. Vizard Agent measures the region first so the zoom lands correctly.

What is a keyword capsule? A small animated word tied to what you just said. Vizard Agent times them to the line.

Can it add a progress bar? Yes, and Vizard Agent verifies it actually advances.

Will it cut my pauses? Yes, if you ask. Dead air in the narration is separate from stalls in the picture.

Can it magnify a specific row? Yes. Vizard Agent measures the bounding box rather than guessing.

What about sensitive data on screen? Tell Vizard Agent, and it will blur rather than magnify.

Does it work for a phone recording? Yes, though the regions are smaller and the crops tighter.

Can it use my own B-roll? Yes. Send it and Vizard Agent prefers yours over anything generated.

How short should an insert be? Short. Vizard Agent treats them as punctuation between actions.

Can it add an end card? Yes. Vizard Agent builds a step review at the end if you ask.

Why cut to the voiceover instead of the recording? Because you clicked in the order the software needed, and you explain in the order a person needs, and those are rarely the same order.