How to cut footage to a narration and keep every shot short
Give Vizard Agent the narration and the footage, and state the maximum shot length as a rule rather than a preference. It transcribes the narration into timed sections, indexes what the footage actually shows, and pins each spoken beat to a span — so the picture follows the words instead of running alongside them.
What is the short version?
The narration is recorded and the footage exists. What you need from Vizard Agent is a cut where every sentence is illustrated by something relevant, and where no single shot outstays its welcome — which on a long narration means several dozen cuts, each one chosen rather than counted out.
- Give Vizard Agent the narration audio and the footage.
- Say the maximum length any single shot may run.
- Say the picture has to match what is being said at that moment.
What do you need before you start?
The narration as audio, and footage with enough coverage to carry it. Vizard Agent can index what each part of the source shows, but if the narration describes five things and the footage only shows three of them, no edit map fixes that.
- The narration. A finished voice track.
- The footage. One long source, or several clips.
- The maximum shot length. Three seconds, four, five.
- The language. Of the narration and any captions.
- Whether the order is fixed. By the narration, usually.
What do you type into Vizard Agent?
State the rule in absolute terms. "Keep it snappy" produces a cut with a few long shots in it; "no shot longer than four seconds, split longer scenes into several cuts" produces a cut you can check against a number.
Prompt
Variants worth knowing:
- "No shot longer than four seconds." A rule, not a mood.
- "Split longer scenes." How to obey the rule without repeating shots.
- "Keep the narration continuous." It sets the length of everything.
What does Vizard Agent actually do?
Here is the Vizard Agent sequence on a narration-led recap cut from one long source video. The part worth reading twice is what happens in the middle, when the cuts start landing on the source's own transitions and the export comes back with white flashes in it.
- Transcribes the narration and surveys the film in parallel.
- Reads both into one compact edit map.
- Groups the narration into timed story chapters.
- Builds a scene index of the source so each beat has candidates to match against.
- Pins the key beats to precise source spans.
- Renders the synced cut under the four-second rule, chapter by chapter.
- Inspects a chunk-render failure and corrects the join list.
- Finds missing duration and repairs overlapping source windows.
- Discovers the cut is using the source's own fade frames as picture.
- Removes the white and black transition ranges from the edit map and re-renders every chapter that used them.
- Extends a chapter left short by the removal, with a valid held frame.
- Trims a stray narration tail and builds a clean ending that does not cut the voice off.
Step nine is the one that separates an automated cut from an edited one. Every dissolve in the source is a run of frames that are half one shot and half another, or simply white — and a cut that lands there looks like a mistake in your video rather than a transition in theirs.
Step ten is the fix, and it is structural rather than cosmetic. Vizard Agent removes those ranges from the map of usable footage, so no chapter can pick them up again, then re-renders the chapters that already had.
Step four is what makes the matching possible at all. Without an index of what the source shows and when, "match the picture to the narration" reduces to cutting on the beat and hoping the coincidence reads as intent.
What does the result look like?
A cut where the picture changes every few seconds, each shot relates to the sentence it sits under, and none of them are the source's dissolves. The narration runs continuously from the first word to the last, and the ending fades rather than stopping mid-breath.
When does this not work well?
Some narration cannot be matched from the footage you have. If the voice track describes things the source never shows, Vizard Agent has the choice of repeating shots or showing something irrelevant, and it will tell you which sections have no coverage rather than quietly padding them.
- Narration beyond the footage. Nothing to show for those lines.
- Very short source material. A rule of four seconds needs a lot of shots.
- Heavily graded transitions. More of the source is unusable than it looks.
- Narration with long pauses. The rule produces cuts over silence.
- Footage that is all one scene. Splitting it makes jump cuts.
How do you fix a result that came back wrong?
Point at one chapter rather than at the whole video. The edit is built as a map of spans rather than as a flat timeline, so Vizard Agent can repin a single beat or exclude one range of source frames and then re-render only the chapters that were affected by the change.
- "There is a white flash at 0:42." That transition range is excluded from the map.
- "This shot does not match the line." That beat is repinned to another span.
- "A shot runs too long." It is split into shorter cuts inside the same span.
- "The voice is cut off at the end." The ending is extended around the narration.
How does Vizard Agent compare to doing it yourself?
By hand this is the long way round: listen, find a shot, trim to four seconds, repeat sixty times. The fades are the part that gets you — you scrub past them, they look fine in the timeline thumbnail, and they show up as a white flash in the export.
| By hand | Vizard Agent | |
|---|---|---|
| Matching | Memory of the footage | An index of what each span shows |
| The length rule | Approximately kept | Applied to every shot |
| Source transitions | Cut into by accident | Excluded from the usable map |
| Repairs | Re-cut the timeline | Re-render the affected chapters |
| The ending | Trimmed to the video | Built around the narration |
Common questions
Does the narration have to be finished? Yes, ideally. Vizard Agent times the whole cut against it.
Can it write the narration too? Vizard Agent can, but a cut built to a voice you recorded yourself is tighter.
What maximum length should I choose? Three to five seconds suits fast-paced work. Vizard Agent will hold to whatever you set.
Can some shots be longer? Yes, if you name the exception — an establishing shot, for example.
What if a section has no matching footage? Vizard Agent tells you rather than filling it with something unrelated.
Does it use every clip I gave it? If you ask for that, yes. Say so, because it constrains the matching.
Why do white flashes appear? They are the source's own dissolves. Vizard Agent strips those ranges out.
Can it add captions? Yes, timed to the narration it already transcribed.
Will the audio be levelled? Yes. The narration is measured and balanced against any other sound.
Can it work from several source videos? Yes. Each one is indexed and both are available for matching.
How long can the finished video be? As long as the narration. Vizard Agent renders in chapters for long ones.
Can I reorder the story? The narration sets the order. Re-record it and Vizard Agent re-cuts to the new one.
What about music underneath? Ask for it, and Vizard Agent mixes it under the narration.
How do I check the length rule was kept? Ask Vizard Agent for the cut list — every shot and its duration.