Vizard Agent

How to cut a long session into clips that each end on a spoken cue

Last updated 2026-08-16 · 6 min read

Give Vizard Agent a long recording and the spoken cue each clip should end on. Vizard Agent transcribes it, extracts the timing of every occurrence of your cue words, plans cuts that keep the whole session in order with nothing skipped, and renders each clip with captions and any overlay you asked for.

What is the short version?

Some recordings have their own structure built into the speech — a word that marks the end of each section. Cutting on that word gives you clips that are genuinely self-contained, in order, with nothing dropped between them.

  1. Go to Vizard Agent and upload the recording.
  2. Say which spoken word each clip should end on.
  3. Say the length range and whether the first clip needs to be longer for an introduction.

What do you need before you start?

The recording and the cue. Vizard Agent finds every instance of the cue word in the transcript, so the requirements you can be precise about — length range, chronological order, nothing skipped — are exactly the ones it can hold to strictly.

What do you type into Vizard Agent?

State the rules as rules. Vizard Agent treats "chronological", "do not skip any footage" and "end each clip on this word" as hard constraints rather than preferences, so listing them plainly is exactly the right way to brief this job.

Prompt

Cut my footage into [60–90] second clips. Keep them chronological, do not skip any footage, make the first about [90] seconds to capture the intro, and end each clip on [the cue word].

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real cue-based split. The middle of the run is spent on the overlay: it built one, tested it on a short clip, looked at the on and off states side by side, and refined it before touching the other nine clips.

  1. Inspects the footage and pulls the transcript and shot list.
  2. Reads the full transcript with all its timestamps and makes a lightweight proxy to work against.
  3. Loads the word timings and samples frames of the speaker.
  4. Extracts the timing of every "yes", "no" and "thank you" in the recording.
  5. Reviews a contact sheet of the speaker's framing across the session.
  6. Builds the glow overlays and plans the exact clip boundaries.
  7. Previews the overlay on a short test clip, then retests it with corrected enable times.
  8. Compares the two overlay states and a plain frame side by side before accepting the look.
  9. Generates captions for all ten clips, re-generates them at a safe vertical position, renders everything and reviews frames from each clip.

Step 8 is the check that matters. A subtle overlay is easy to make too strong or too weak, and Vizard Agent judged it by putting the states next to each other rather than by looking at one of them in isolation.

What does the result look like?

From the run this page is written from, probed on a delivered clip: 1600x900, H.264, 30fps, 53.07 seconds, AAC audio — one of ten clips cut in a single pass, in chronological order with the whole session covered and each piece ending on its cue.

Vizard Agent verified the durations and reviewed frames from the middle and end of every clip before uploading them, so the whole set was checked rather than sampled.

Ten clips out of one session is a month of posts if you space them out. They arrive as separate files with their own links, which is what makes scheduling them practical rather than a second editing job.

Timing was not measured separately for this kind of job. The closest measured work runs a median of 28 to 38 minutes end to end, with the middle half spread considerably wider. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.

When does this not work well?

Cue-based cutting is exact, and exactness cuts both ways: if the cue word appears somewhere you did not intend, the clip ends there. Vizard Agent follows the rule you gave it rather than deciding when the rule was probably not meant.

How do you fix a result that came back wrong?

Name the clip number. Vizard Agent keeps the transcript, every cue timing, the overlay tests and the cut plan, so adjusting a boundary or restyling the captions across all ten is a re-render from work it already did rather than a fresh analysis.

How does Vizard Agent compare to doing it yourself?

By hand this is listening for a word across a long session, noting every occurrence, then cutting ten clips and captioning each — and every rule you set for yourself has to be enforced by memory as the evening wears on.

By hand Vizard Agent
Finding the cue Listen and note Every occurrence timed from the transcript
Holding to the rules Remember them Applied as constraints
Testing an overlay Render and judge Tested on a clip, states compared
Ten clips Ten export cycles Rendered and reviewed as a set

Common questions

What makes a good cue word? Something you say deliberately at the end of a section. A common conversational word will fire in the wrong places.

Can it keep everything in order? Yes. Chronological order with nothing skipped is a constraint Vizard Agent holds to exactly.

Can the first clip be longer? Yes. Say so and Vizard Agent keeps the introduction whole in the opening piece rather than splitting it across two clips.

Can it add captions to all of them? Yes, and cheaply, because the word timings come from the same transcription.

Can it add a visual effect tied to what is said? Yes. Vizard Agent builds it, tests it on a short clip and compares the states before applying it everywhere.

How many clips will I get? As many as the cue and the length range produce. Vizard Agent reports the durations when it delivers them.

Can it reframe the clips for a feed? Yes. Vizard Agent re-frames on the speaker rather than cropping the centre of a wide shot.

Does it work in any language? Yes. Vizard Agent transcribes in the language spoken and finds the cue word in it.