Vizard Agent

How to pick clips on how someone sounds rather than what they said

Last updated 2026-09-17 · 8 min read

Tell Vizard Agent which qualities of delivery you are after — breathlessness, effort, a moment where someone loses their thread — and it reviews the audio for those rather than searching the transcript for good lines. A transcript is a record of words, and the thing that makes a clip land is often audible in how they were said.

What is the short version?

Searching a transcript finds you the well-phrased sentences, and a well-phrased sentence delivered comfortably is very often the least interesting thirty seconds in a thirty-minute recording. What makes somebody stop scrolling is hearing a person who is obviously working hard.

  1. Tell Vizard Agent which delivery qualities you want.
  2. Ask it to review the audio, not only the transcript.
  3. Ask it to treat stumbles as candidates, not faults.

What do you need before you start?

The recording and a sense of what you want the clip to feel like. Vizard Agent works from a described quality — winded, emotional, flustered, certain — so naming the feeling gives it something to search for that a transcript cannot provide.

What do you type into Vizard Agent?

Describe the sound to Vizard Agent rather than the subject. Asking for "the motivational bits" returns the scripted lines every time; asking for "where she is genuinely out of breath" returns the moments that actually read as real to a viewer.

Prompt

This is a thirty-minute workout shot on a locked camera. Do not just pick the motivational lines — go through the audio and find the sections where she is genuinely working, breathing hard and sounding tired, plus any moment where she stumbles or loses her thread. Those are the Shorts I want.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on a real single-take workout recording cut into vertical Shorts. Note that a transcript still gets built along the way — it is simply used to navigate the recording rather than to choose the moments.

  1. Samples frames across the full recording to see the framing.
  2. Transcribes the session to locate speech, stumbles and pauses.
  3. Summarises the transcript rather than reading it end to end.
  4. Reads the remaining spoken moments for coherent segments.
  5. Views representative frames from the spoken and high-effort sections.
  6. Checks the audio-review tool for verifying breathing and delivery.
  7. Reviews the long candid section for fatigue, heavy breathing and an arc.
  8. Reviews the motivational section for clarity and fatigue in the delivery.
  9. Checks a late stumble as a natural, relatable moment.
  10. Builds the chosen Shorts with captions timed to the real words.

Step six is the one that changes the outcome. Reading a transcript tells you somebody said "keep going"; listening tells you whether they said it easily or forced it out, and only the second version is worth posting.

Step nine is a deliberate inversion. A stumble is normally something you cut around, and in a piece whose appeal is effort it is the most honest few seconds in the recording.

Step three is a practical habit worth copying. A thirty-minute transcript read in full is expensive and mostly filler, so it is summarised first and only the promising windows are read properly.

What does the result look like?

Clips that sound like somebody doing the thing rather than presenting it. The breathing is in them, the effort is audible, and Vizard Agent tells you which quality it selected each clip for so you can argue with the criterion rather than the result.

When does this not work well?

Delivery is not always the selling point of a piece. A technical explainer, a legal update or a product walkthrough are all worth clipping for what is actually said in them, and sending Vizard Agent hunting for emotion produces clips that are lively and completely uninformative.

How do you fix a result that came back wrong?

Say what you heard rather than simply which clip you disliked. Notes like "that one sounds performed" and "that stumble is just a mistake" are both criteria in their own right, and Vizard Agent can apply either of them across the rest of the recording.

How does Vizard Agent compare to doing it yourself?

By hand you search the transcript, because the transcript is searchable and the audio is thirty minutes long that nobody wants to sit through twice. So you find the sentences and miss the moments entirely, and the clips come out competent, well-phrased and completely flat.

By hand Vizard Agent
What gets searched The transcript The audio, guided by the transcript
What gets found Well-phrased lines Moments of real effort
Stumbles Cut around Considered as candidates
Reading the transcript End to end, or not at all Summarised, then read in windows
The criterion Implicit Named, and reported per clip

Common questions

Why not just use the transcript? Because it records words, not delivery. Vizard Agent uses it to navigate.

What qualities can it look for? Effort, fatigue, hesitation, emotion. Describe it and Vizard Agent searches for it.

Is a stumble really usable? Often, in anything whose appeal is authenticity. Vizard Agent will offer them.

Can it tell a stumble from an error? With guidance. Tell Vizard Agent which ones you rejected and why.

Does this work for interviews? Yes, especially where the emotion matters more than the phrasing.

What about a locked camera? Fine. The selection is by audio; the framing is a separate problem.

Will it still caption them? Yes, timed to the real words rather than to the script.

What if the audio is poor? Breathing is the first thing lost. Vizard Agent will say when it cannot hear it.

How many clips should I ask for? Say the number and Vizard Agent sets the bar to match.

Can I mix both criteria? Yes. Ask for some on content and some on delivery.

Does it need the whole recording? Yes. Anything Vizard Agent did not hear cannot be among the moments it offers.

Can it explain its choices? Yes. Vizard Agent reports which quality it picked each individual clip for.

What about music underneath? Music masks breathing entirely, so Vizard Agent makes its selections before any is added.

Will the clips overlap? Not unless you want them to. Say so and Vizard Agent keeps each one to its own window of the recording.