Vizard Agent

How to pick the one podcast clip a specific audience needs

Last updated 2026-08-25 · 7 min read

Upload the interview and describe who the clip is for. Vizard Agent transcribes everything, searches for the language that audience would recognise in themselves, checks the framing at every candidate cut point, tracks where the speakers' heads sit, and cuts one clip rather than a handful.

What is the short version?

Most podcast clipping selects for what is most quotable. Selecting for a particular audience is a different question — what would make this specific group of people stop scrolling — and it usually produces a different passage from the same hour.

  1. Go to Vizard Agent and upload the full interview.
  2. Describe the audience in detail, not just the topic.
  3. Say you want one clip, and how long it should be.

What do you need before you start?

The interview and a real description of the reader. Vizard Agent searches the transcript for the language that group uses about itself, so a few sentences about who they are and what they are quietly dealing with is worth more than a topic keyword.

What do you type into Vizard Agent?

Describe the audience the way you would describe a person. Vizard Agent uses that description to search the transcript for matching language, so "successful women litigators whose precision has become a burden" gives it far more to find than "lawyers".

Prompt

Create ONE [LinkedIn] video clip from this full podcast interview. Audience: [who they are, what drives them, what it costs them]. Around [60] seconds, [4:5], captions.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real single-clip selection from a full interview. The steps that matter most are two: searching the transcript for the audience's own language, and measuring where the faces sit before deciding on the crop.

  1. Checks the interview file and the project history, then transcribes the whole thing.
  2. Samples frames from the interview and looks at the footage.
  3. Finds the key passages in the transcript and reads each candidate section.
  4. Searches for the specific language the audience would recognise — the words they use about themselves.
  5. Reads further context around each candidate before committing to one.
  6. Checks the framing at every planned cut point and looks at those shots.
  7. Extracts word-level cut points and gets exact timings for the remaining sections.
  8. Detects the face position for the vertical crop and measures where it sits in frame.
  9. Tracks the second speaker's cap to find head position and shot changes, then crops and captions the clip.

Step 4 is the difference between a clip and the right clip. Searching for the language of identity loss found a passage that a search for the episode's stated topic would have walked straight past.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1350, H.264, 24fps, 60 seconds, AAC audio. Four-by-five portrait, exactly a minute, one passage lifted from a full interview with both speakers kept in frame and captions running throughout.

Four-by-five is the shape a professional feed gives the most room to, and twenty-four frames a second keeps it looking like a filmed conversation rather than a screen recording. Both are choices, not defaults.

When does this not work well?

Selecting one clip for one audience is a judgement call, and it depends entirely on the passage genuinely existing somewhere in the recording. These are the cases where Vizard Agent runs out of material rather than out of care.

How do you fix a result that came back wrong?

Say what the clip missed or got wrong. Vizard Agent keeps the full transcript, every candidate passage it considered, the word-level cut points and the face measurements, so a different passage can be cut without re-reading the whole interview.

How does Vizard Agent compare to doing it yourself?

By hand this means listening to the whole interview with a particular reader in mind, which is genuinely hard to sustain for an hour, and then cropping a two-person conversation into a vertical frame without losing either face.

By hand Vizard Agent
Selection Listen for something quotable Search for the audience's own language
Coverage Whatever you remember Every candidate passage read
Cut points Guess the sentence start Word-level boundaries
Cropping Centre and hope Faces detected and tracked

Common questions

Why one clip instead of several? Because selecting for a specific audience gives Vizard Agent one strongest answer, not five equal ones.

How detailed should the audience description be? Very. Vizard Agent searches the transcript for language that matches how they describe themselves.

Can it keep both speakers in frame? Yes. Vizard Agent measures where each head sits before choosing the crop.

What if the interview is very long? That is fine. Vizard Agent transcribes and reads all of it through before selecting anything.

Will the clip start cleanly? Yes. Vizard Agent uses word-level cut points rather than rounding to the second.

Why search for language rather than topic? Because people do not scroll past a clip because it is on their topic; they stop because someone said the thing they have been thinking. That phrasing is what the transcript search is looking for.

Can I get alternatives? Yes. Ask and Vizard Agent cuts the runner-up passages it had already identified and read.

What shape suits a professional feed? Four-by-five portrait, which Vizard Agent used here. It takes more space than a square without going full vertical.

Does it caption automatically? Yes. Vizard Agent captions from the word timings it already extracted.

Can I use the same interview for other audiences? Yes, and you should. Vizard Agent already holds the transcript, so cutting a second clip for a different group costs a fraction of the first one.

How long should a professional clip run? Around a minute. Vizard Agent has room for one complete thought at that length, and a professional feed rarely holds attention for a longer single-idea clip than that.

Does the audience description have to be a real segment? It has to be specific enough to have its own vocabulary. Vizard Agent is matching phrasing in the transcript, so a group defined only by job title gives it far less to work with than one defined by what they are actually wrestling with.

Does Vizard Agent check the finished clip? Yes. It reviews the framing and the caption timing before delivery.