How to cut a long session into clips that each end on a spoken cue
Give Vizard Agent a long recording and the spoken cue each clip should end on. Vizard Agent transcribes it, extracts the timing of every occurrence of your cue words, plans cuts that keep the whole session in order with nothing skipped, and renders each clip with captions and any overlay you asked for.
What is the short version?
Some recordings have their own structure built into the speech — a word that marks the end of each section. Cutting on that word gives you clips that are genuinely self-contained, in order, with nothing dropped between them.
- Go to Vizard Agent and upload the recording.
- Say which spoken word each clip should end on.
- Say the length range and whether the first clip needs to be longer for an introduction.
What do you need before you start?
The recording and the cue. Vizard Agent finds every instance of the cue word in the transcript, so the requirements you can be precise about — length range, chronological order, nothing skipped — are exactly the ones it can hold to strictly.
- The recording. One long session. Length is not a constraint.
- The cue word. The word or phrase that marks the end of a section.
- The length range. Sixty to ninety seconds is a normal window for this kind of clip.
- A longer first clip. Say so if the introduction needs to fit in the opening piece.
- Any overlay. A colour glow, a badge, a caption style tied to what was said.
What do you type into Vizard Agent?
State the rules as rules. Vizard Agent treats "chronological", "do not skip any footage" and "end each clip on this word" as hard constraints rather than preferences, so listing them plainly is exactly the right way to brief this job.
Prompt
Variants worth knowing:
- A visual cue too. Ask for a subtle colour treatment tied to a spoken answer, as in the run below.
- Captions on every clip. The word timings already exist, so it is a small addition.
- A different cue per section. Several cue words can be handled in one pass.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real cue-based split. The middle of the run is spent on the overlay: it built one, tested it on a short clip, looked at the on and off states side by side, and refined it before touching the other nine clips.
- Inspects the footage and pulls the transcript and shot list.
- Reads the full transcript with all its timestamps and makes a lightweight proxy to work against.
- Loads the word timings and samples frames of the speaker.
- Extracts the timing of every "yes", "no" and "thank you" in the recording.
- Reviews a contact sheet of the speaker's framing across the session.
- Builds the glow overlays and plans the exact clip boundaries.
- Previews the overlay on a short test clip, then retests it with corrected enable times.
- Compares the two overlay states and a plain frame side by side before accepting the look.
- Generates captions for all ten clips, re-generates them at a safe vertical position, renders everything and reviews frames from each clip.
Step 8 is the check that matters. A subtle overlay is easy to make too strong or too weak, and Vizard Agent judged it by putting the states next to each other rather than by looking at one of them in isolation.
What does the result look like?
From the run this page is written from, probed on a delivered clip: 1600x900, H.264, 30fps, 53.07 seconds, AAC audio — one of ten clips cut in a single pass, in chronological order with the whole session covered and each piece ending on its cue.
Vizard Agent verified the durations and reviewed frames from the middle and end of every clip before uploading them, so the whole set was checked rather than sampled.
Ten clips out of one session is a month of posts if you space them out. They arrive as separate files with their own links, which is what makes scheduling them practical rather than a second editing job.
Timing was not measured separately for this kind of job. The closest measured work runs a median of 28 to 38 minutes end to end, with the middle half spread considerably wider. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.
When does this not work well?
Cue-based cutting is exact, and exactness cuts both ways: if the cue word appears somewhere you did not intend, the clip ends there. Vizard Agent follows the rule you gave it rather than deciding when the rule was probably not meant.
- A common cue word appears everywhere. "Yes" in casual conversation will produce boundaries you did not want.
- Fixed lengths and cue words can conflict. If no cue falls inside the window, something has to give — say which.
- Covering everything means keeping the dull parts. "Do not skip any footage" is a real instruction with real consequences.
- Poor audio misses cues. A quiet or overlapped cue word may not appear in the transcript at all.
- Ten clips is ten posts to schedule. The cutting is the easy half of that.
How do you fix a result that came back wrong?
Name the clip number. Vizard Agent keeps the transcript, every cue timing, the overlay tests and the cut plan, so adjusting a boundary or restyling the captions across all ten is a re-render from work it already did rather than a fresh analysis.
- "Clip four ends too early." The next cue instead of the first.
- "The glow is too strong." Refined and compared side by side again.
- "Move the captions higher." Regenerated at a safer position across the set.
How does Vizard Agent compare to doing it yourself?
By hand this is listening for a word across a long session, noting every occurrence, then cutting ten clips and captioning each — and every rule you set for yourself has to be enforced by memory as the evening wears on.
| By hand | Vizard Agent | |
|---|---|---|
| Finding the cue | Listen and note | Every occurrence timed from the transcript |
| Holding to the rules | Remember them | Applied as constraints |
| Testing an overlay | Render and judge | Tested on a clip, states compared |
| Ten clips | Ten export cycles | Rendered and reviewed as a set |
Common questions
What makes a good cue word? Something you say deliberately at the end of a section. A common conversational word will fire in the wrong places.
Can it keep everything in order? Yes. Chronological order with nothing skipped is a constraint Vizard Agent holds to exactly.
Can the first clip be longer? Yes. Say so and Vizard Agent keeps the introduction whole in the opening piece rather than splitting it across two clips.
Can it add captions to all of them? Yes, and cheaply, because the word timings come from the same transcription.
Can it add a visual effect tied to what is said? Yes. Vizard Agent builds it, tests it on a short clip and compares the states before applying it everywhere.
How many clips will I get? As many as the cue and the length range produce. Vizard Agent reports the durations when it delivers them.
Can it reframe the clips for a feed? Yes. Vizard Agent re-frames on the speaker rather than cropping the centre of a wide shot.
Does it work in any language? Yes. Vizard Agent transcribes in the language spoken and finds the cue word in it.