Vizard Agent

How to cut a face-off clip that stays on one person

Last updated 2026-08-30 · 7 min read

Upload the footage and tell Vizard Agent whose clip this is. Vizard Agent finds the stare-down, the escalation and the moment it is broken up, keeps the other party on screen only enough to explain the confrontation, and builds the tension around one person's face.

What is the short version?

A confrontation involves two people and a clip about it usually does not. Following one of them turns a piece of coverage into a character study, and that means choosing shots by whose expression they carry rather than by what is happening.

  1. Go to Vizard Agent and upload the footage.
  2. Say which person the clip follows.
  3. Say where it starts — the stare-down, the first word, the entrance.

What do you need before you start?

The footage and a subject. Vizard Agent locates the beats itself, so timestamps are unnecessary, but naming who the clip belongs to changes almost every shot choice it makes and is the one thing it will not guess.

What do you type into Vizard Agent?

Say who it follows and why. Vizard Agent selects for presence rather than for action when a subject is named, which means it will hold on a face during a pause instead of cutting to whoever is speaking.

Prompt

Cut a high-energy vertical clip from this [face-off]. Focus almost entirely on [name] — expressions, confidence, presence. Start on the stare-down. Show [the other party] only enough to set up the confrontation. Build tension through to [the intervention].

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real face-off clip. The detail worth noticing comes late: an emoji in the on-screen text would not render, so it found a typeface that supported it and tested that before rebuilding the video.

  1. Locates the uploaded footage and checks the available opening effects.
  2. Measures the video and extracts a visual sample of the main moments.
  3. Brings the footage into the project and builds a contact sheet.
  4. Reads the contact sheet to identify the stare-down, the escalation and the intervention.
  5. Analyses the audio and the action at the start of the confrontation and at the flashpoint.
  6. Extracts closer frames of both moments and reviews the subject's expressions up close.
  7. Pulls exact timings for the speech and the pauses so cuts do not break the audio.
  8. Builds the first vertical cut with subtle reframing and on-screen text, then renders a review sheet and waveform.
  9. Checks the composition and text legibility at full size, finds a typeface that renders the emoji, tests it as an image, then picks a better closing frame of the subject and rebuilds.

Step 9 is two small saves in one. An emoji that renders as an empty box ruins a caption, and a clip that ends wherever the source ended feels unfinished — both are only visible if you look at the actual frames.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1072x1920, H.264, 30fps, 66.37 seconds, AAC audio. Vertical, just over a minute, opening on the stare-down and building through the escalation to the moment it is broken up, with the subject on screen for almost all of it.

The other party appears, but only to establish who is being faced. That imbalance is deliberate and it is what makes the clip belong to one person rather than being neutral coverage of an event.

When does this not work well?

Following one person through a confrontation depends entirely on the camera having stayed on them, and on you having the right to use the footage in the first place. Vizard Agent can only cut what was filmed and what you are entitled to publish.

How do you fix a result that came back wrong?

Name the beat that is wrong and say what it should be instead. Vizard Agent keeps the contact sheets, the close frames, the speech timings and the render review sheets, so the opening or the ending can change without recutting the middle.

How does Vizard Agent compare to doing it yourself?

By hand this means watching a press conference several times to find the three beats that matter, then cropping a wide shot into vertical without cutting anyone's head off. The emoji that does not render is discovered after upload.

By hand Vizard Agent
Finding the beats Watch it repeatedly Located from a contact sheet and audio
Expressions Scrub for the good ones Close frames pulled and reviewed
Cuts Trim by ear Cut against speech and pause timings
Text Discover the problem live Checked at full size before delivery

Common questions

Can Vizard Agent follow one person? Yes. Name them and it selects for their presence rather than for whoever is speaking.

How does it find the flashpoint? Vizard Agent analyses the audio and the action, then reviews close frames of both moments.

Will the cuts break the audio? No. Vizard Agent uses speech and pause timings so cuts land between phrases.

Can it reframe a wide shot to vertical? Yes, subtly. Vizard Agent checks at full size that the faces survive the crop.

Can I add text over it? Yes. Vizard Agent verifies it does not cover faces and that every character renders.

Why go back for a different closing frame? Because a clip that stops where the footage stops feels like it ran out rather than ended. Choosing the last frame deliberately is what makes it read as a finished piece.

How long should a clip like this be? Around a minute. Vizard Agent builds the tension across it and resolves on the flashpoint.

Do I need the rights to the footage? Yes, and for televised events those usually sit with the broadcaster, not with you.

Can Vizard Agent add commentary? Vizard Agent can, though on-screen text usually works better for a tension-based clip.

Can I cut a version following the other person? Yes, and it is nearly free. Vizard Agent already has the contact sheets and the timings, so the second clip is a re-selection against a different subject rather than a fresh pass over the footage.

Does Vizard Agent check the finished clip? Yes. It reviews the composition, the text and the last frame before delivery.