How to stop a screen recording stalling under the voiceover
Send the screen recording and the separately recorded voiceover, and give Vizard Agent two rules: show whatever the narration is talking about, and never let the picture sit still for more than a few seconds. It finds the stalls and fills them with keyword cards, magnified detail or related shots.
What is the short version?
A tutorial built from narration plus a screen capture has one very specific failure mode: the voice keeps explaining while the screen sits on a static settings panel. Nothing about it is technically wrong, every word is correct, and the viewer leaves anyway.
- Go to Vizard Agent with the recording and the voiceover.
- Say the picture must show what the narration is describing.
- Say how long the screen may sit still before something must happen.
What do you need before you start?
Both files, and a number. The number is the important part — "don't let it get boring" is not actionable, and "never more than five to eight seconds without visible movement" is a rule that can be applied and checked across a whole video.
- The screen recording. As captured.
- The voiceover. Recorded separately.
- The stall limit. Five to eight seconds is a good default.
- What may fill a stall. B-roll, keyword cards, icons, zooms.
- How long a filler may run. Short; it is punctuation, not content.
What do you type into Vizard Agent?
Give the rule and the exception together. The instruction that carries the whole job is not "make it engaging" — it is that when the screen stops moving, something else has to, and you say how long the picture is allowed to wait first.
Prompt
Variants worth knowing:
- "Show what I'm talking about." The sync rule.
- "Never still past 5 seconds." The stall rule, as a number.
- "Keep the inserts short." Stops filler becoming the video.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real product tutorial built from a screen capture and a separate narration. Steps two and three are what makes the rest possible: it has to know exactly when each action happens.
- Checks both files and loads the tutorial editing approach.
- Analyses the narration words against the screen picture in parallel.
- Samples frames densely to locate the precise moment of each action.
- Detects the window and editor boundaries in the capture.
- Measures the crop boundaries for each region of the interface.
- Renders the segments against the narration's line, not the recording's order.
- Checks each segment's opening frame for composition.
- Renders a progress bar and section titles over the top.
- Builds the opening title and keyword animations, and lays in the voiceover.
- Splits a section into a wide shot and a close-up where one shot was not enough.
- Analyses the assembled cut for sync and flow.
- Adds an end card reviewing the steps.
- Diagnoses a mis-crop and a missing element found on review.
- Builds a magnifier effect aligned to one specific line on screen.
- Measures the preview area's exact bounding box so the magnifier lands.
- Re-renders the magnifier until it sits on the right line.
- Adds keyword capsules timed to the narration.
- Adds a slow push-in to the section that was otherwise static.
- Diagnoses a repeated second and a jump cut and rebuilds those.
- Verifies the progress bar advances and the ending holds.
Step six is the reordering that people do not expect. The screen recording happened in the order you clicked; the narration explains things in the order that makes sense — so the picture is cut to the voiceover's sequence rather than played back as recorded.
Steps fourteen to sixteen are what a stall gets filled with when there is nothing else to show. The screen was static on one panel, so a magnifier was built over the exact line being discussed — measured against the real bounding box, because a magnifier over the wrong row is worse than no magnifier.
What does the result look like?
A tutorial where the picture is always doing something, and always doing the specific thing being described at that moment. Where the interface genuinely had nothing left to show, Vizard Agent carries those seconds with a keyword card, a slow push-in or a magnified detail instead.
The inserts are short by design. A stall filled with a ten-second stock clip is a different problem wearing the same clothes — the fix is punctuation, not a second video running underneath the first.
When does this not work well?
The method needs the recording to actually contain the thing the narration is describing. Where the voiceover talks at length about a screen that was never captured, no amount of pacing work by Vizard Agent invents that screen, and the honest answer is to record it.
- Narration about unrecorded screens. Nothing to cut to.
- Very long explanations of one panel. Filler cannot carry a minute.
- A voiceover that does not match the flow. Re-record or re-script.
- Low-resolution captures. Zooms and magnifiers fall apart.
- Sensitive data on screen. Say so; magnifying it makes it worse.
How do you fix a result that came back wrong?
Point at the second where it goes wrong. Vizard Agent keeps the measured action timings, the interface region boundaries and the narration alignment from the first pass, so a mis-timed insert or a wrong crop becomes a rebuild of that one segment rather than of the whole tutorial.
- "That is the wrong panel." Re-cropped from the measured boundaries.
- "The magnifier is on the wrong line." Re-measured and re-aligned.
- "Too many inserts." The stall limit relaxed.
- "It repeats a second." That join diagnosed and rebuilt.
How does Vizard Agent compare to doing it yourself?
By hand this is the most tedious kind of editing there is: scrub to find where you clicked, match it to where you said it, and then decide what to do about the forty seconds where you explain something while the screen shows a settings panel.
| By hand | Vizard Agent | |
|---|---|---|
| Finding each action | Scrubbing | Located by dense sampling |
| Matching to the narration | By ear, clip by clip | Cut to the voiceover's order |
| Stalls | Noticed on playback | Found against a stated limit |
| Filling them | Whatever is to hand | Keyword cards, zooms, magnifiers |
| Checking | Watch it through | Sync and flow analysed |
Common questions
Do I record the voiceover first? Either way. Vizard Agent cuts the picture to the voice, so voice-first is easier.
What is a good stall limit? Five to eight seconds. Vizard Agent applies whatever number you give it.
Will it reorder my recording? Yes, to follow the narration. Say so if the order must be preserved.
Can it zoom instead of cutting away? Yes. Vizard Agent measures the region first so the zoom lands correctly.
What is a keyword capsule? A small animated word tied to what you just said. Vizard Agent times them to the line.
Can it add a progress bar? Yes, and Vizard Agent verifies it actually advances.
Will it cut my pauses? Yes, if you ask. Dead air in the narration is separate from stalls in the picture.
Can it magnify a specific row? Yes. Vizard Agent measures the bounding box rather than guessing.
What about sensitive data on screen? Tell Vizard Agent, and it will blur rather than magnify.
Does it work for a phone recording? Yes, though the regions are smaller and the crops tighter.
Can it use my own B-roll? Yes. Send it and Vizard Agent prefers yours over anything generated.
How short should an insert be? Short. Vizard Agent treats them as punctuation between actions.
Can it add an end card? Yes. Vizard Agent builds a step review at the end if you ask.
Why cut to the voiceover instead of the recording? Because you clicked in the order the software needed, and you explain in the order a person needs, and those are rarely the same order.