Vizard Agent

How to pull several Reels out of one long video link

Last updated 2026-08-24 · 7 min read

Send Vizard Agent the link and ask for short Reels. Vizard Agent downloads the video, transcribes it, finds the passages that stand alone, computes where the speaker's face sits in each one, tests the crop before committing, and tracks the framing when the speaker moves.

What is the short version?

A long video contains a handful of moments that work on their own, and the rest is context. Getting them out means finding passages that make sense without setup, then solving the crop, because a widescreen speaker does not fit a vertical frame by accident.

  1. Go to Vizard Agent and send the link.
  2. Say how many Reels you want and roughly how long.
  3. Say vertical, and whether you want captions.

What do you need before you start?

A link. Vizard Agent downloads the video, reads it and picks the passages itself, so there is nothing to prepare and no timecodes to mark. Saying what the Reels are for helps, since a clip that works as a teaser is a different passage from one that works as advice.

What do you type into Vizard Agent?

One sentence is enough. Vizard Agent reads the whole transcript and decides which passages stand alone, so you do not have to specify the moments, and naming a topic only helps if you already know which part you want.

Prompt

Can you pull some short Reels out of this video: [link]. [Two] of them, vertical, with captions.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real repurposing job. The reframing work is more than half of it, because turning a widescreen conversation into a vertical Reel is a per-clip problem rather than a single crop setting.

  1. Downloads the video from the link and probes its duration and format.
  2. Transcribes the whole thing and extracts frames from the significant moments.
  3. Reviews the candidate clips as frame sheets and compares them.
  4. Reads the word-level timings and inspects each candidate passage's text and time range.
  5. Computes the face positions and crop coordinates for each clip rather than centring blindly.
  6. Tests several crop angles, builds the trials as images and reviews them side by side.
  7. Detects the scene changes inside a clip and checks the wide-angle frames separately.
  8. Builds a tracking crop for the clip where the speaker moves and tests the camera transitions.
  9. Generates re-timed captions for each Reel independently, previews the subtitle rendering on frames, builds and rebuilds the title graphics, then renders both Reels at full quality and verifies their levels.

Step 5 is the difference between a repurposed clip and a broken one. Centring a vertical crop on the middle of a widescreen frame puts the speaker at the edge as often as not, and computing where the face actually is costs one measurement per clip.

What does the result look like?

From the run this page is written from, probed on one delivered Reel: 1080x1920, H.264, 30fps, 54.3 seconds, AAC audio. Vertical, just under a minute, reframed onto the speaker with a tracking crop, animated captions and a title graphic on the front.

Two Reels from one video is a realistic yield rather than a disappointing one. Most long videos contain two or three genuinely standalone passages, and cutting six would mean four of them needing context the viewer does not have.

When does this not work well?

Repurposing depends on two things being true: the source has genuinely self-contained moments in it, and the original framing can be recovered into a vertical crop. Vizard Agent finds and reframes whatever is there, and some videos simply resist both of those at once.

How do you fix a result that came back wrong?

Name the clip that needs changing. Vizard Agent keeps the full transcript, the word timings, the face measurements, every crop trial and both rendered Reels, so re-cutting a passage or adjusting a crop is a re-render against work Vizard Agent has already done.

How does Vizard Agent compare to doing it yourself?

By hand this means scrubbing a long video for passages you half remember, then discovering that a straight vertical crop puts the speaker's ear at the centre of the frame. Keyframing a crop that follows someone across a shot is the part that turns a ten-minute job into an afternoon.

By hand Vizard Agent
Finding standalone passages Scrub and guess Read from the full transcript
The vertical crop Centre it and hope Face positions computed per clip
A speaker who moves Keyframe by hand Tracking crop built and tested
Captions per clip Retype or re-sync Re-timed independently for each Reel

Common questions

Is a link enough? Yes. Vizard Agent downloads the video from the link and works from the file directly.

How many Reels can it get? Usually two or three from a single video. Push beyond that and the clips start needing context the viewer does not have.

Will the speaker stay in frame? Yes. Vizard Agent computes the face position and builds a tracking crop where the speaker moves.

Does each Reel get its own captions? Yes. Vizard Agent re-times them independently for each clip rather than slicing one caption track into pieces, which is what keeps the timing correct after the cut.

Why not just crop to the centre? Because in a widescreen conversation the speaker is rarely centred. A blind vertical crop puts them at the edge or cuts them in half, which is why Vizard Agent measures the face position in every candidate before choosing a crop.

Does Vizard Agent test the crop before rendering? Yes. It builds the crop trials as images and reviews them side by side, so the framing is chosen by looking rather than by assuming the speaker is centred.

Can it add a hook card? Yes. Vizard Agent builds a title graphic for the front of each Reel and checks it on frame.

What if the original has burnt-in graphics? They come along with the crop. Vizard Agent can frame around a lower third where there is room in the shot, and a graphic sitting over the speaker cannot be avoided.

Does Vizard Agent check the finished Reels? Yes. It pulls quality-control frames from each and verifies the audio levels before uploading.