Vizard Agent

How to turn a long Zoom call into a short highlight reel

Last updated 2026-08-23 · 6 min read

Upload the recording and tell Vizard Agent how many takeaways you want. Vizard Agent transcribes the call, groups it into speaker turns to find where the real answers are, checks alternative soundbites for each point, and builds chapter cards so the reel reads as a summary rather than a set of offcuts.

What is the short version?

An hour-long call contains perhaps six minutes worth keeping. Finding them means reading the whole transcript as a conversation — who asked what, who answered at length — rather than scanning for keywords, because the substance sits inside the long answers.

  1. Go to Vizard Agent and upload the recording.
  2. Say how many takeaways you want and how long the reel should be.
  3. Say whether you want chapter cards and captions.

What do you need before you start?

The recording and a number. Vizard Agent reads the whole call and decides what the takeaways are, so you do not need to mark anything, and telling it how many points and how long the result should be is what shapes the selection it makes.

What do you type into Vizard Agent?

Ask for takeaways rather than highlights. Vizard Agent reads "the five best takeaways from the points discussed" as an instruction to find substance in the answers, whereas "highlights" invites it to pick lively moments that do not add up to a summary.

Prompt

Take this [60] minute call and create a [3] minute highlight edit of the [5] best takeaways from the points discussed.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real call. The selection work is at the top and the audio work is at the bottom, and both halves ended with Vizard Agent checking its own result and redoing part of it.

  1. Probes the recording and extracts any embedded subtitles from it.
  2. Transcribes the audio with word timings and checks the transcript structure.
  3. Extracts sample frames across the call and reviews the layout as it changes.
  4. Analyses the full transcript across the timeline and groups it into speaker turns.
  5. Scans the long answers to find where the substance actually is.
  6. Extracts the speech segments for the five takeaways and verifies the word timestamps for every candidate clip.
  7. Checks alternative soundbites for the same points before committing.
  8. Looks at the framing of each candidate moment and measures the duration and loudness of every extracted clip.
  9. Generates intro, chapter and summary cards, tests a lower-third on a real frame, renders the cut, measures speech against music on the finished file and re-renders with the dialogue level restored.

Step 7 is the one that raises the quality. The first clip that says a thing is rarely the best clip that says it, and checking the alternatives costs a search of a transcript Vizard Agent already holds.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1280x720, H.264, 25fps, 165.52 seconds, AAC audio. Two minutes and forty-six seconds, five takeaways with chapter cards, burned subtitles, a lower-third and a music bed under the speech.

Two minutes forty-six against a three-minute brief is Vizard Agent letting the content decide. Padding a summary to hit a round number adds material that was rejected for a reason, so the reel ends when the fifth takeaway does.

When does this not work well?

A call recording is a document of a private conversation before it is anything else, and Vizard Agent treats it purely as footage to be read and cut. Everything about who agreed to what, and what may be shared outside the call, stays outside its judgement.

How do you fix a result that came back wrong?

Name the takeaway. Vizard Agent keeps the transcript, the speaker turns, the verified word timings, the alternative soundbites and the rendered cards, so swapping a point is a re-cut against work it already did across the full hour.

How does Vizard Agent compare to doing it yourself?

By hand this is watching a full hour twice over, noting timecodes for half a dozen candidate moments, cutting them together, building the chapter cards separately, and then finding that the dialogue has been quietly attenuated by the music bed.

By hand Vizard Agent
Finding the substance Scan for keywords Grouped into speaker turns, answers read
Choosing a soundbite Take the first one Alternatives checked before committing
Chapter cards Build them separately Generated and reviewed on frame
The dialogue level Trust the export Measured on the render, then corrected

Common questions

Does this work for other long recordings? Yes. The method transfers to any recording where the value is in what people said — a webinar, a panel, a board meeting, a customer interview. Vizard Agent reads the transcript as a conversation rather than scanning it for keywords, which is what surfaces the substantial answers.

How long a call can it handle? An hour is routine here. Vizard Agent transcribes the whole recording and reads all of it rather than sampling.

Does it know who said what? Vizard Agent groups the transcript into speaker turns, which is how it locates the long answers where the substance sits.

Will it add captions? Yes, burned in and aligned to the timeline, generated from the same word timings Vizard Agent used to cut the clips.

Can I ask for three points instead of five? Yes. Say the number and the selection narrows accordingly.

Does it use the embedded Zoom subtitles? It checks them and transcribes independently for word-level timings.

Can I get a one-minute version? Yes. Ask for both and the short cut comes from the same takeaways.

Does Vizard Agent check the finished mix? Yes. It measures the speech and the music separately on the render, which is how it caught the dialogue being attenuated here and re-rendered.

Will it hit the exact length I asked for? Close to it. Vizard Agent lets the content end the reel rather than padding it.