Vizard Agent

How to compile every clip of one speaker from a long recording

Last updated 2026-08-16 · 6 min read

Give Vizard Agent a long recording and say whose parts you want. Vizard Agent transcribes it with speaker labels, checks the frames to confirm the right person is on screen at those timestamps, and builds the compilation two ways — one that keeps the natural flow, one that keeps only that speaker's words.

What is the short version?

Somebody spoke for eleven minutes across a two-hour panel, and you want those eleven minutes. Finding them is a transcript problem; cutting them cleanly is a seeking problem, and getting the audio to survive the cut is a third problem entirely.

  1. Go to Vizard Agent and give it the recording or a link to it.
  2. Name the speaker you want kept.
  3. Say whether the cut should flow naturally or include only their words.

What do you need before you start?

The recording and a name. Vizard Agent separates the speakers itself and checks the picture to confirm the labelling, so what you supply is which of them matters and how strictly you want the compilation to hold to them.

What do you type into Vizard Agent?

Name the person and the recording. Vizard Agent handles the speaker separation and the verification, so you do not need to supply timestamps or say how many sections there are — it works both out from the transcript.

Prompt

Compile all the parts where [name] is speaking into one video. [link]

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real speaker compilation. The first half finds the sections and the whole second half is verification — and that verification caught two separate faults in the render that an ordinary export would have shipped without anyone noticing until playback.

  1. Downloads the recording from the link.
  2. Generates a word-timed transcript with speaker labels.
  3. Extracts frames at the timestamps where that speaker is talking.
  4. Looks at the frame grid to confirm the right person is on screen at those times.
  5. Inspects the word-level file and analyses the exact word boundaries for that speaker.
  6. Builds both versions — natural flow and strict — and uploads them for analysis.
  7. Analyses both renders for cut cleanliness, finds visual glitches, and re-renders with frame-accurate seeking.
  8. Checks the codec streams and the volume, and discovers the audio track has come out silent.
  9. Works out that the order of the seek flags is the cause, re-renders both versions correctly, and verifies the audio is present and loud.

Steps 8 and 9 are the reason this page exists. A silent audio track is invisible in a progress log and obvious to the first person who watches the file; Vizard Agent measured the volume rather than assuming it was there.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1920x1080, H.264, 25fps, 97.46 seconds, AAC audio. Wide, just over a minute and a half of one speaker cut from a much longer recording, with clean joins and verified audio.

Both versions are delivered, so you can compare a cut that flows naturally against one that contains nothing but that speaker's words.

It is not instant, and this category was not separately measured. Comparable work runs a median of 28 to 38 minutes end to end. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.

When does this not work well?

Speaker separation is good and not perfect, and a compilation is only ever as reliable as the labelling underneath it. Vizard Agent checks the picture against the transcript at the timestamps it plans to cut, and some recordings are genuinely ambiguous about who is talking.

How do you fix a result that came back wrong?

Say what is wrong with which section. Vizard Agent keeps the transcript, the speaker boundaries, both cut lists and the verified renders, so dropping a section or loosening the cut is a re-render from timings it already measured.

How does Vizard Agent compare to doing it yourself?

By hand this is scrubbing a two-hour recording with a notepad, writing down every in and out point, cutting, and then discovering the export is silent or the cuts glitch — which is when the evening actually starts.

By hand Vizard Agent
Finding their sections Scrub and note Speaker-labelled transcript
Confirming it is them Trust your ear Frames checked at those timestamps
Clean cuts Notice a glitch later Analysed, then re-rendered accurately
Audio surviving Find out on playback Volume measured and verified

Common questions

Does it know who is who? It separates the speakers and checks the frames at those timestamps to confirm the labelling matches the picture.

Can it handle a two-hour recording? Yes. Length is not the constraint. Vizard Agent transcribes the whole thing and works from the timings, so a long panel costs time rather than accuracy.

What is the difference between the two versions? One keeps the natural flow around their answers; the other keeps only their words. Both are delivered.

Can it work from a link? Yes. Vizard Agent downloads the recording from the link and works from the file, so you do not need to upload a large panel video yourself.

Can I get vertical clips instead of one compilation? Yes. Ask for their strongest sections as separate clips, re-framed for a feed.

Will the audio be right? Vizard Agent measures the volume of the finished render rather than assuming the track came through.

Can it add captions? Yes. The word timings already exist from the transcript, so captioning the compilation is a small extra step.

Can it cut by topic instead? Yes. Name the subject and Vizard Agent searches the transcript for it rather than for a speaker.

How long will the compilation be? However long they spoke. Vizard Agent does not pad or trim to a target unless you ask for one.

Can it drop a section I do not want? Yes. Name it and Vizard Agent re-renders without it from the timings it already has.