Vizard Agent

How to cut footage to a narration and keep every shot short

Last updated 2026-09-17 · 8 min read

Give Vizard Agent the narration and the footage, and state the maximum shot length as a rule rather than a preference. It transcribes the narration into timed sections, indexes what the footage actually shows, and pins each spoken beat to a span — so the picture follows the words instead of running alongside them.

What is the short version?

The narration is recorded and the footage exists. What you need from Vizard Agent is a cut where every sentence is illustrated by something relevant, and where no single shot outstays its welcome — which on a long narration means several dozen cuts, each one chosen rather than counted out.

  1. Give Vizard Agent the narration audio and the footage.
  2. Say the maximum length any single shot may run.
  3. Say the picture has to match what is being said at that moment.

What do you need before you start?

The narration as audio, and footage with enough coverage to carry it. Vizard Agent can index what each part of the source shows, but if the narration describes five things and the footage only shows three of them, no edit map fixes that.

What do you type into Vizard Agent?

State the rule in absolute terms. "Keep it snappy" produces a cut with a few long shots in it; "no shot longer than four seconds, split longer scenes into several cuts" produces a cut you can check against a number.

Prompt

Sync this footage to the voiceover track. Cut the picture to match what the narration is saying at each point. Strict rule: no shot longer than four seconds — split longer scenes into several shorter cuts that still match the words. Keep the narration continuous from start to finish.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on a narration-led recap cut from one long source video. The part worth reading twice is what happens in the middle, when the cuts start landing on the source's own transitions and the export comes back with white flashes in it.

  1. Transcribes the narration and surveys the film in parallel.
  2. Reads both into one compact edit map.
  3. Groups the narration into timed story chapters.
  4. Builds a scene index of the source so each beat has candidates to match against.
  5. Pins the key beats to precise source spans.
  6. Renders the synced cut under the four-second rule, chapter by chapter.
  7. Inspects a chunk-render failure and corrects the join list.
  8. Finds missing duration and repairs overlapping source windows.
  9. Discovers the cut is using the source's own fade frames as picture.
  10. Removes the white and black transition ranges from the edit map and re-renders every chapter that used them.
  11. Extends a chapter left short by the removal, with a valid held frame.
  12. Trims a stray narration tail and builds a clean ending that does not cut the voice off.

Step nine is the one that separates an automated cut from an edited one. Every dissolve in the source is a run of frames that are half one shot and half another, or simply white — and a cut that lands there looks like a mistake in your video rather than a transition in theirs.

Step ten is the fix, and it is structural rather than cosmetic. Vizard Agent removes those ranges from the map of usable footage, so no chapter can pick them up again, then re-renders the chapters that already had.

Step four is what makes the matching possible at all. Without an index of what the source shows and when, "match the picture to the narration" reduces to cutting on the beat and hoping the coincidence reads as intent.

What does the result look like?

A cut where the picture changes every few seconds, each shot relates to the sentence it sits under, and none of them are the source's dissolves. The narration runs continuously from the first word to the last, and the ending fades rather than stopping mid-breath.

When does this not work well?

Some narration cannot be matched from the footage you have. If the voice track describes things the source never shows, Vizard Agent has the choice of repeating shots or showing something irrelevant, and it will tell you which sections have no coverage rather than quietly padding them.

How do you fix a result that came back wrong?

Point at one chapter rather than at the whole video. The edit is built as a map of spans rather than as a flat timeline, so Vizard Agent can repin a single beat or exclude one range of source frames and then re-render only the chapters that were affected by the change.

How does Vizard Agent compare to doing it yourself?

By hand this is the long way round: listen, find a shot, trim to four seconds, repeat sixty times. The fades are the part that gets you — you scrub past them, they look fine in the timeline thumbnail, and they show up as a white flash in the export.

By hand Vizard Agent
Matching Memory of the footage An index of what each span shows
The length rule Approximately kept Applied to every shot
Source transitions Cut into by accident Excluded from the usable map
Repairs Re-cut the timeline Re-render the affected chapters
The ending Trimmed to the video Built around the narration

Common questions

Does the narration have to be finished? Yes, ideally. Vizard Agent times the whole cut against it.

Can it write the narration too? Vizard Agent can, but a cut built to a voice you recorded yourself is tighter.

What maximum length should I choose? Three to five seconds suits fast-paced work. Vizard Agent will hold to whatever you set.

Can some shots be longer? Yes, if you name the exception — an establishing shot, for example.

What if a section has no matching footage? Vizard Agent tells you rather than filling it with something unrelated.

Does it use every clip I gave it? If you ask for that, yes. Say so, because it constrains the matching.

Why do white flashes appear? They are the source's own dissolves. Vizard Agent strips those ranges out.

Can it add captions? Yes, timed to the narration it already transcribed.

Will the audio be levelled? Yes. The narration is measured and balanced against any other sound.

Can it work from several source videos? Yes. Each one is indexed and both are available for matching.

How long can the finished video be? As long as the narration. Vizard Agent renders in chapters for long ones.

Can I reorder the story? The narration sets the order. Re-record it and Vizard Agent re-cuts to the new one.

What about music underneath? Ask for it, and Vizard Agent mixes it under the narration.

How do I check the length rule was kept? Ask Vizard Agent for the cut list — every shot and its duration.