Vizard Agent

How to pull out only the parts of a long video about one topic

Last updated 2026-08-16 · 6 min read

Upload the long recording to Vizard Agent and describe the topic you want kept. Vizard Agent transcribes the whole thing, filters the transcript for segments about that subject, reads the lines around each match to find where the passage really begins and ends, then joins them with crossfades into one video.

What is the short version?

You have forty minutes of a class, a talk or a session, and the eight minutes you actually want are scattered through it. Finding those eight minutes means listening to all forty with a notepad, which is why the shorter version never gets made.

  1. Go to Vizard Agent and upload the recording.
  2. Describe the topic in your own words — the ideas, not the keywords.
  3. Say whether you want one continuous video or separate clips.

What do you need before you start?

The recording and a description of the theme. Vizard Agent searches the meaning of the transcript rather than matching literal words, so you can describe what you are after the way you would explain it to a colleague, and it will find the passages where the speaker circles the same idea in different language.

What do you type into Vizard Agent?

Describe the theme in plain sentences. Vizard Agent reads the transcript for meaning, so a paragraph explaining what you want kept will find more of the right material than a list of terms, and it will not miss passages that make your point without ever using your words.

Prompt

This is a [41] minute [yoga class]. Make a video with only the parts where she talks about [relaxing the mind and slowing down]. Keep it flowing naturally.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real topic extraction. The stretch in the middle is the part that matters: having found each matching segment, it reads the lines on either side of it, because a passage almost never starts exactly where the keyword lands.

  1. Finds the uploaded file and analyses the video's format and streams.
  2. Checks how the transcription tool works and what it costs before running it.
  3. Transcribes the full forty-one minutes.
  4. Inspects the transcript's structure and a sample segment to understand its shape.
  5. Filters the segments that are about the requested theme.
  6. Reads the surrounding segments at every hit — seven separate passages, each examined in context.
  7. Extracts frames from several points and looks at them to judge the composition.
  8. Cuts all seven passages concurrently rather than one after another.
  9. Calculates the crossfade offsets, merges them, and extracts frames at each transition to check the joins look smooth.

Step 6 is what separates this from a keyword search. Vizard Agent looked at the lines before and after each match to find where the speaker actually starts and finishes the thought, which is the difference between a passage that makes sense and one that begins mid-sentence.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1920x1080, H.264, 30fps, 425.97 seconds, AAC audio. Seven minutes cut out of forty-one, in the recording's own shape and resolution, with crossfades between the passages rather than hard joins.

Vizard Agent checked those transitions by extracting frames at each dissolve and looking at them, so the joins were verified visually rather than assumed to be fine.

Timing was not measured separately for this kind of job. The closest measured work runs a median of 28 to 38 minutes end to end, with the middle half spread considerably wider. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.

When does this not work well?

The transcript is the whole search index, so anything the speaker did not say out loud is invisible to this method. Vizard Agent reads meaning rather than keywords and reads around every match for context, and it still cannot find a moment that exists only in the picture.

How do you fix a result that came back wrong?

Name the passage. Vizard Agent keeps the full transcript, the segment boundaries it chose and the cut pieces, so dropping a section or widening one to include its setup is a re-render against timings it already has rather than a second transcription of the whole recording.

How does Vizard Agent compare to doing it yourself?

By hand this is watching forty-one minutes with a stopwatch, writing down in and out points, then cutting and hoping you did not miss a passage in the twelve minutes where your attention drifted. Doing it properly means watching it twice.

By hand Vizard Agent
Finding the passages Watch it all, twice Transcript filtered by meaning
Setting the boundaries Guess and adjust Surrounding lines read at each hit
Cutting seven pieces One at a time Extracted concurrently
Checking the joins Watch the render Frames pulled at every dissolve

Common questions

Does it search for exact words? No. Vizard Agent reads the transcript for meaning, so it finds passages that make your point in different words.

How long can the source be? An hour is routine. Vizard Agent transcribes the whole recording once and works from the timings, so a longer source costs time rather than accuracy.

Will the cuts sound abrupt? Vizard Agent reads around each match to find natural boundaries and joins the pieces with crossfades rather than hard cuts.

Can I get the passages separately? Yes. Say so and each one is delivered as its own file instead of a single compilation.

Can it do more than one topic? Yes, and it is cheaper together, because the transcription only happens once.

Can Vizard Agent add captions to the extract? Yes. The word timings already exist from the transcription, so captioning the finished compilation is a small extra step.

Will it change the resolution? No. Vizard Agent delivers in the source's own shape unless you ask for something else.