How to compile every clip of one speaker from a long recording
Give Vizard Agent a long recording and say whose parts you want. Vizard Agent transcribes it with speaker labels, checks the frames to confirm the right person is on screen at those timestamps, and builds the compilation two ways — one that keeps the natural flow, one that keeps only that speaker's words.
What is the short version?
Somebody spoke for eleven minutes across a two-hour panel, and you want those eleven minutes. Finding them is a transcript problem; cutting them cleanly is a seeking problem, and getting the audio to survive the cut is a third problem entirely.
- Go to Vizard Agent and give it the recording or a link to it.
- Name the speaker you want kept.
- Say whether the cut should flow naturally or include only their words.
What do you need before you start?
The recording and a name. Vizard Agent separates the speakers itself and checks the picture to confirm the labelling, so what you supply is which of them matters and how strictly you want the compilation to hold to them.
- The recording. A file or a link. Length is not a problem.
- Which speaker. By name, or "the one on the left".
- How strict. Natural flow keeps a little context around each answer; strict keeps only their words.
- The output shape. Wide for a talk, vertical if it is going to a feed.
- Whether you need both versions. They cost little more together than one.
What do you type into Vizard Agent?
Name the person and the recording. Vizard Agent handles the speaker separation and the verification, so you do not need to supply timestamps or say how many sections there are — it works both out from the transcript.
Prompt
Variants worth knowing:
- Both versions. Ask for the natural-flow cut and the strict cut; comparing them tells you which you actually wanted.
- A topic as well as a speaker. "Only where they talk about X" narrows it further.
- Vertical clips instead. Ask for their strongest sections as separate short clips.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real speaker compilation. The first half finds the sections and the whole second half is verification — and that verification caught two separate faults in the render that an ordinary export would have shipped without anyone noticing until playback.
- Downloads the recording from the link.
- Generates a word-timed transcript with speaker labels.
- Extracts frames at the timestamps where that speaker is talking.
- Looks at the frame grid to confirm the right person is on screen at those times.
- Inspects the word-level file and analyses the exact word boundaries for that speaker.
- Builds both versions — natural flow and strict — and uploads them for analysis.
- Analyses both renders for cut cleanliness, finds visual glitches, and re-renders with frame-accurate seeking.
- Checks the codec streams and the volume, and discovers the audio track has come out silent.
- Works out that the order of the seek flags is the cause, re-renders both versions correctly, and verifies the audio is present and loud.
Steps 8 and 9 are the reason this page exists. A silent audio track is invisible in a progress log and obvious to the first person who watches the file; Vizard Agent measured the volume rather than assuming it was there.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1920x1080, H.264, 25fps, 97.46 seconds, AAC audio. Wide, just over a minute and a half of one speaker cut from a much longer recording, with clean joins and verified audio.
Both versions are delivered, so you can compare a cut that flows naturally against one that contains nothing but that speaker's words.
It is not instant, and this category was not separately measured. Comparable work runs a median of 28 to 38 minutes end to end. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.
When does this not work well?
Speaker separation is good and not perfect, and a compilation is only ever as reliable as the labelling underneath it. Vizard Agent checks the picture against the transcript at the timestamps it plans to cut, and some recordings are genuinely ambiguous about who is talking.
- Overlapping speech confuses the split. Two people talking at once belongs to both or neither.
- Similar voices get merged. Vizard Agent verifies against the frames, and off-camera speakers have no picture to check.
- Strict cuts lose the question. An answer without its question often makes no sense. That is what the natural-flow version is for.
- Quotes taken out of context are your responsibility. Compiling someone's remarks changes what surrounds them.
- Publishing someone else's recording is a separate question. Cutting it is not the same as having the right to post it.
How do you fix a result that came back wrong?
Say what is wrong with which section. Vizard Agent keeps the transcript, the speaker boundaries, both cut lists and the verified renders, so dropping a section or loosening the cut is a re-render from timings it already measured.
- "Keep the question before each answer." The natural-flow version, or wider boundaries.
- "That section is someone else." The boundary corrected against the frames.
- "Make it vertical." Re-framed on the speaker rather than cropped.
How does Vizard Agent compare to doing it yourself?
By hand this is scrubbing a two-hour recording with a notepad, writing down every in and out point, cutting, and then discovering the export is silent or the cuts glitch — which is when the evening actually starts.
| By hand | Vizard Agent | |
|---|---|---|
| Finding their sections | Scrub and note | Speaker-labelled transcript |
| Confirming it is them | Trust your ear | Frames checked at those timestamps |
| Clean cuts | Notice a glitch later | Analysed, then re-rendered accurately |
| Audio surviving | Find out on playback | Volume measured and verified |
Common questions
Does it know who is who? It separates the speakers and checks the frames at those timestamps to confirm the labelling matches the picture.
Can it handle a two-hour recording? Yes. Length is not the constraint. Vizard Agent transcribes the whole thing and works from the timings, so a long panel costs time rather than accuracy.
What is the difference between the two versions? One keeps the natural flow around their answers; the other keeps only their words. Both are delivered.
Can it work from a link? Yes. Vizard Agent downloads the recording from the link and works from the file, so you do not need to upload a large panel video yourself.
Can I get vertical clips instead of one compilation? Yes. Ask for their strongest sections as separate clips, re-framed for a feed.
Will the audio be right? Vizard Agent measures the volume of the finished render rather than assuming the track came through.
Can it add captions? Yes. The word timings already exist from the transcript, so captioning the compilation is a small extra step.
Can it cut by topic instead? Yes. Name the subject and Vizard Agent searches the transcript for it rather than for a speaker.
How long will the compilation be? However long they spoke. Vizard Agent does not pad or trim to a target unless you ask for one.
Can it drop a section I do not want? Yes. Name it and Vizard Agent re-renders without it from the timings it already has.