Vizard Agent

How to cut gaming highlights that keep your facecam in shot

Last updated 2026-08-23 · 6 min read

Upload the session and tell Vizard Agent to keep both your face and the gameplay visible. Vizard Agent reads the frames and the transcript together, measures the audio peaks to find your reactions, checks each moment at the frame before and the frame during, and cuts without splitting your sentences.

What is the short version?

The best moment in a horror stream is not the jump scare, it is your face half a second later. Finding those means looking at the reaction as well as the game, and keeping both on screen means never cropping one away to fit the other.

  1. Go to Vizard Agent and upload the session recording.
  2. Say the highlights must show both your face and the gameplay.
  3. Say roughly how many moments and how long.

What do you need before you start?

The recording and the rule. Vizard Agent finds the moments by watching the game and listening to you at the same time, so nothing needs marking, and stating that both the face and the gameplay must stay visible is the constraint that shapes every crop decision.

What do you type into Vizard Agent?

State the constraint plainly. Vizard Agent treats "without cutting away my face" as a layout rule rather than a preference, which stops it choosing crops that would give a cleaner gameplay frame at the cost of your reaction.

Prompt

Pull the best moments out of my session without ever cutting away my face or the gameplay. Around [10] moments.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real session recording. The reason the method works is that Vizard Agent kept opening the reaction frames rather than the scare itself — your face in the moment immediately after it happened.

  1. Probes the recording and extracts frames across the whole session.
  2. Looks at a frame from every few minutes — the start, the third minute, the sixth, and so on to the end.
  3. Transcribes what you say during the game.
  4. Searches specifically for the scares, the chases and your strongest reactions.
  5. Analyses the session in three passes — the first ten minutes, the middle, the last stretch.
  6. Measures the audio peaks and extracts frames at each candidate moment.
  7. Looks at each moment closely — the instant before the scare, the peak, and your reaction to it.
  8. Cuts the eleven best moments and reviews them all together as a grid.
  9. Joins the clips, adds the title, reviews the cut frame by frame and the waveform, then fixes a cut that was splitting a sentence and re-checks the join.

Step 9's fix is the one that separates a reel from a rough export. A clip that ends mid-sentence sounds broken even when the moment inside it is great, and catching that means listening to the assembly rather than trusting the timestamps.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1920x1080, H.264, 60fps, 101.57 seconds, AAC audio. Widescreen at sixty frames per second, a minute and forty seconds, eleven moments cut from a half-hour session with a title over the opening.

Keeping sixty frames per second matters for gameplay in a way it does not for most video. Halving it would smooth out exactly the fast motion the highlights exist to show, so Vizard Agent delivered at the source rate.

When does this not work well?

Vizard Agent finds the moments from what is actually visible on screen and audible on the track. A quiet session, or a facecam scaled down into a corner, limits how much there is for it to work with in the first place.

How do you fix a result that came back wrong?

Name the moment. Vizard Agent keeps the transcript, the audio peak analysis, every extracted frame and the cut clips, so dropping a moment or extending one is a re-cut against work it already did across the whole session.

How does Vizard Agent compare to doing it yourself?

By hand this is scrubbing half an hour of footage for the moments you half remember, trimming each of them down, and then watching the reel back to find that three start mid-word and one cuts you off completely.

By hand Vizard Agent
Finding moments Scrub for them Frames, transcript and audio peaks together
The reaction Take the scare Reaction frames checked separately
Cut points Trim to the action Verified against the sentence
Layout Crop for the game Face and gameplay both kept visible

Common questions

Why not just clip the loudest moments? Because volume finds the scare and not the clip. The audio peak tells Vizard Agent where to look, and the frame a beat later — your reaction — is what decides whether the moment is worth cutting at all.

Will my face always be visible? Yes, if you say so. Vizard Agent treats it as a layout rule rather than a preference.

How long a session can it handle? Half an hour was routine on this run, and Vizard Agent reads all of it rather than sampling.

Does it keep sixty frames per second? Yes. Gameplay needs it, so the delivery matches the source rate.

Can I get a vertical version? Yes, with the facecam stacked over the game rather than one cropped away.

Why check the reaction frames separately? Because the moment worth clipping is usually a beat after the moment in the game. Looking only at the gameplay finds the scare; looking at your face finds the clip.

Can Vizard Agent add subtitles? Yes, built from the same transcript it already made in order to find the moments in the first place.

Does Vizard Agent check the finished reel? Yes. It reviews the assembled cut frame by frame and inspects the waveform, which is how it caught a clip splitting a sentence on this run.

Does it work for other genres? Yes. Horror, racing, shooters — what changes is what counts as a moment, not the method.