Vizard Agent

How to cut several vertical highlight clips with subtitles

Last updated 2026-08-16 · 6 min read

Give Vizard Agent a long video and say how many vertical clips you want. Vizard Agent transcribes it, finds the strongest moments, tests the 9:16 crop on real frames to confirm the subject is not cut off, and burns subtitles that fill the frame edge to edge with no black bars.

What is the short version?

Cutting highlights out of a long video is two problems wearing one coat: finding the moments worth keeping, and reframing wide footage to vertical without decapitating whoever is speaking. The second one is where automatic tools usually fail.

  1. Go to Vizard Agent and upload the long video.
  2. Say how many clips you want and roughly how long each should be.
  3. Say the subtitle language and that the frame must be filled with no bars.

What do you need before you start?

The video and a count. Vizard Agent finds the moments itself from the transcript and the picture, so the useful instructions are about the output — how many clips, what language the subtitles are in, and whether anything in the frame must stay visible.

What do you type into Vizard Agent?

Say the count, the shape and the subtitle language in one line. Vizard Agent chooses the moments from what is actually said and shown, so you do not need to supply timestamps unless there is a specific section you want covered.

Prompt

Cut [5] vertical highlight clips from this video. Each with [language] subtitles, filled to 9:16 with no black bars at the top or bottom.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real highlights batch. The crop is not assumed to be safe — it renders a test frame at the target shape and looks at it before any clip is committed to that framing.

  1. Loads the repurposing playbook and probes the video's streams and duration.
  2. Transcribes the audio with timestamps.
  3. Extracts frames at several points and looks at them to understand the composition.
  4. Lists the asset capabilities and reads the reframing tool's options.
  5. Builds a shot map of the video to find precise cut points.
  6. Renders test crops at 9:16 and inspects them to confirm the subject stays in frame.
  7. Writes a script to slice the word-level transcript and shift each clip's timestamps to start at zero.
  8. Generates the subtitle files for all five clips in parallel rather than one at a time.
  9. Lists the installed font families and finds the exact font file path so the burn-in renders correctly.

Steps 6 and 9 are the two that fail silently otherwise. A centre crop that beheads the speaker looks fine in a progress log, and a missing font draws empty boxes instead of throwing an error — so Vizard Agent looked at both with its own eyes.

What does the result look like?

From the run this page is written from, probed on a delivered clip: 540x960, H.264, 25fps, 17.0 seconds, AAC audio. Vertical and filled edge to edge, seventeen seconds, with burned-in subtitles — one of five clips cut from the same source in a single pass.

Ask for the clips at a higher resolution and Vizard Agent delivers larger files, though the source's own resolution is the ceiling on what any of them can be.

This category was not separately measured, so treat the timing as a range rather than a promise: comparable work runs a median of 28 to 38 minutes end to end. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.

When does this not work well?

Reframing wide footage to vertical throws away most of the picture, and no amount of care changes that arithmetic. Vizard Agent tests the crop and follows the subject rather than the centre, and some compositions simply do not survive the journey.

How do you fix a result that came back wrong?

Say which clip. Vizard Agent keeps the transcript, the shot map, the tested crops and each clip's subtitle file separately, so re-framing one clip or extending its ending does not disturb the other four or repeat the transcription.

How does Vizard Agent compare to doing it yourself?

By hand this is watching the whole video with a notepad, cutting five sections, reframing each one by dragging a crop box to follow the speaker, then typing and timing subtitles five separate times. Every stage is simple and every stage has to be repeated once per clip.

By hand Vizard Agent
Finding moments Watch the whole thing Transcript plus shot map
Reframing Drag the crop box Test frame rendered and checked
Subtitles Type and time each clip Sliced from word timings, in parallel
Fonts Find out it broke on render Font file located first

Common questions

How many clips can it cut at once? Five is a normal batch and Vizard Agent generates their subtitles in parallel rather than one after another.

Will the subject stay in frame? Vizard Agent tests the crop on real frames and follows the subject rather than cropping the centre blindly.

Can the subtitles be in a different language from the speech? Yes. Say which language you want and Vizard Agent generates them accordingly.

Will there be black bars? Not if you say so. The frame is filled by cropping rather than by shrinking the wide picture.

Can I pick the sections myself? Yes. Give timestamps or describe the part and Vizard Agent works within it.

Why does a batch take a while? Transcription, shot mapping and five separate burns are real work. It is minutes rather than seconds.

Can it add a headline to each clip? Yes. Vizard Agent sets one over the opening frames, which matters on feeds where the first frame sells the clip.

Does it work on a podcast recording? Yes. Two-person framing is the case to watch — say which speaker should be followed.