Vizard Agent

How to put a highlights recap in front of the full video

Last updated 2026-09-28 · 8 min read

Ask Vizard Agent for a cold open of the strongest moments, a short pause, then the complete video with only the silences shortened. The rule that makes it work is that no spoken word is removed — so the cutting is driven by measured gaps between words rather than by a target runtime.

What is the short version?

You have a long talk that is worth watching in full, and nobody will start it without a reason. A recap at the front supplies the reason, and the body still contains everything you said. Vizard Agent builds both from one set of word timings.

  1. Tell Vizard Agent to open with a recap, then run the full video.
  2. Say that only pauses may be cut, never speech.
  3. Expect some pauses to be protected because they belong to the sentence.

What do you need before you start?

The full recording and a decision about what the recap is for. A recap that previews the best lines is a different thing from one that teases without revealing, and the first is what suits a video someone is about to watch anyway.

Be explicit that the body stays complete. "Make it shorter" and "make it shorter without losing any content" produce very different edits, and only the second one leaves you with the full talk.

What do you type into Vizard Agent?

State the structure and the prohibition in the same message to Vizard Agent. The structure is the easy half to describe; the prohibition is the part that stops a well-meant edit from quietly dropping a sentence somewhere in the middle in order to hit a length target.

Put the most interesting and intense parts in the first seconds as a summary, then a pause, then the full video — but cut the useless pauses so it runs as short as possible. Do not remove any spoken content, only the pauses.

That last sentence is doing the real work. Without it, "as short as possible" is an invitation to trim content, which is the opposite of what the structure is for.

What does Vizard Agent actually do?

It transcribes word by word before it cuts anything, precisely so that every word can be protected. The silences between words are then measured, and the cut list is built from those measurements rather than from a sense of where the video drags.

In that session the passes ran:

Then the discovery that changed the rule: a review pass flagged a visual jump, and the cause was that pauses inside sentences had been cut. Those were protected and the silence rule refined, which kept the natural hesitation in the delivery.

What does the result look like?

A video that opens on its best two or three moments, breathes for a beat, and then plays through in full — noticeably shorter than the original recording while still containing every single word that was spoken in it.

The recap ends properly. Its final extract was extended so it closes on a complete sentence rather than being cut off mid-thought, and that change was made to the intro alone without re-rendering the body.

When does this not work well?

When the talk has no real peaks in it. A steady, evenly paced explanation has nothing that stands out enough to build a recap from, and a montage of three fairly ordinary moments ends up promising something the video does not then deliver.

Removing silence also has a floor. A speaker who already talks briskly has little dead air, so the body will barely shorten — and pushing past that point means cutting speech, which is the thing you ruled out.

And the structure costs you a clean start. Anyone who watches from the beginning sees the best moments twice, which is fine for a video people choose to watch and wrong for one that plays automatically.

How do you fix a result that came back wrong?

If the speech sounds clipped or rushed anywhere, the silence rule has gone too far. Ask Vizard Agent to protect the pauses that fall inside sentences — that single change is usually the whole difference between an edit that feels tight and one that feels unnatural.

If you see a jump at a cut, give the timestamp. A visual jump almost always marks a removed pause that was carrying a gesture or a breath, and Vizard Agent restores that gap rather than smoothing the picture over it.

If the recap ends abruptly, ask for its last extract to be extended to a complete sentence. Vizard Agent can rebuild the intro without touching the body, so that fix is cheap.

How does Vizard Agent compare to doing it yourself?

Silence removal exists in every editor and most of them cut on a threshold, which is exactly the behaviour that produces the jumps. The tool has no idea whether a gap is dead air or a speaker pausing for effect, and a threshold cannot tell those apart.

Working from word timings can. Knowing where each word starts and ends is what allows a cut to be placed between sentences rather than inside one, and it is also what makes the promise — no spoken content removed — something you can actually verify rather than hope for.

Common questions

Will any of my words be cut? No, if you say so. Vizard Agent cuts only measured silence between words.

How much shorter will it get? However much dead air there is. Vizard Agent reports the figure rather than targeting one.

Why protect pauses inside sentences? Because they are part of the delivery. Cutting them reads as a jump and sounds unnatural.

How long should the recap be? A few seconds of the strongest moments. Vizard Agent picks them from the transcript peaks.

Does the recap need captions? It helps, since the moments arrive without context. Ask Vizard Agent to add them.

Should there be a pause after the recap? Yes. Vizard Agent uses it to separate the summary from the real start.

Can the recap end mid-sentence? It should not. Vizard Agent extends the last extract to a complete sentence.

Can I change the recap later? Yes, without re-rendering the body. Vizard Agent rebuilds the intro alone.

Does the audio stay in sync? Yes. Vizard Agent checks the picture and sound alignment after the silence cuts.

What if the recap repeats too much? Say so and Vizard Agent shortens it or picks moments from further apart.

Can it choose moments visually as well? Yes. Vizard Agent builds a shot map alongside the transcript for that.

Will the delivered file be checked? Yes. Vizard Agent checks the uploaded copy rather than the local render.

Does this work on a podcast? Yes, and Vizard Agent finds it a good fit: long, spoken, worth hearing in full.

Can it do the reverse, a teaser that reveals nothing? Yes, though that is a different brief and Vizard Agent picks different moments.