How to cut gaming highlights that keep your facecam in shot
Upload the session and tell Vizard Agent to keep both your face and the gameplay visible. Vizard Agent reads the frames and the transcript together, measures the audio peaks to find your reactions, checks each moment at the frame before and the frame during, and cuts without splitting your sentences.
What is the short version?
The best moment in a horror stream is not the jump scare, it is your face half a second later. Finding those means looking at the reaction as well as the game, and keeping both on screen means never cropping one away to fit the other.
- Go to Vizard Agent and upload the session recording.
- Say the highlights must show both your face and the gameplay.
- Say roughly how many moments and how long.
What do you need before you start?
The recording and the rule. Vizard Agent finds the moments by watching the game and listening to you at the same time, so nothing needs marking, and stating that both the face and the gameplay must stay visible is the constraint that shapes every crop decision.
- The session. Half an hour is normal. Vizard Agent reads all of it.
- The rule. Face and gameplay both visible.
- How many moments. Ten or so makes a good reel.
- The shape. Widescreen keeps both; vertical means a stacked layout.
- A title. Vizard Agent will put one over the opening.
What do you type into Vizard Agent?
State the constraint plainly. Vizard Agent treats "without cutting away my face" as a layout rule rather than a preference, which stops it choosing crops that would give a cleaner gameplay frame at the cost of your reaction.
Prompt
Variants worth knowing:
- Vertical. Say so and Vizard Agent stacks the facecam over the game.
- Subtitles. Worth it; a lot of gaming content is watched muted.
- A specific kind of moment. Scares, fails, wins — it narrows the search.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real session recording. The reason the method works is that Vizard Agent kept opening the reaction frames rather than the scare itself — your face in the moment immediately after it happened.
- Probes the recording and extracts frames across the whole session.
- Looks at a frame from every few minutes — the start, the third minute, the sixth, and so on to the end.
- Transcribes what you say during the game.
- Searches specifically for the scares, the chases and your strongest reactions.
- Analyses the session in three passes — the first ten minutes, the middle, the last stretch.
- Measures the audio peaks and extracts frames at each candidate moment.
- Looks at each moment closely — the instant before the scare, the peak, and your reaction to it.
- Cuts the eleven best moments and reviews them all together as a grid.
- Joins the clips, adds the title, reviews the cut frame by frame and the waveform, then fixes a cut that was splitting a sentence and re-checks the join.
Step 9's fix is the one that separates a reel from a rough export. A clip that ends mid-sentence sounds broken even when the moment inside it is great, and catching that means listening to the assembly rather than trusting the timestamps.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1920x1080, H.264, 60fps, 101.57 seconds, AAC audio. Widescreen at sixty frames per second, a minute and forty seconds, eleven moments cut from a half-hour session with a title over the opening.
Keeping sixty frames per second matters for gameplay in a way it does not for most video. Halving it would smooth out exactly the fast motion the highlights exist to show, so Vizard Agent delivered at the source rate.
When does this not work well?
Vizard Agent finds the moments from what is actually visible on screen and audible on the track. A quiet session, or a facecam scaled down into a corner, limits how much there is for it to work with in the first place.
- A tiny facecam is hard to read. Your reaction has to be legible for it to be found.
- Vertical forces a compromise. Both elements fit stacked, and each gets less room.
- Not every session has eleven moments. Ask for fewer if the session was quiet.
- Streamed audio may include others. Anyone else in the call is in the reel too.
- Game content has platform rules. Some titles and some moments are restricted where you post.
How do you fix a result that came back wrong?
Name the moment. Vizard Agent keeps the transcript, the audio peak analysis, every extracted frame and the cut clips, so dropping a moment or extending one is a re-cut against work it already did across the whole session.
- "The ouija bit is better than the basement one." Swapped from the candidates it found.
- "That clip cuts me off." Re-cut at the sentence, as this run did.
- "Make it vertical." Re-laid out with the facecam stacked over the game.
How does Vizard Agent compare to doing it yourself?
By hand this is scrubbing half an hour of footage for the moments you half remember, trimming each of them down, and then watching the reel back to find that three start mid-word and one cuts you off completely.
| By hand | Vizard Agent | |
|---|---|---|
| Finding moments | Scrub for them | Frames, transcript and audio peaks together |
| The reaction | Take the scare | Reaction frames checked separately |
| Cut points | Trim to the action | Verified against the sentence |
| Layout | Crop for the game | Face and gameplay both kept visible |
Common questions
Why not just clip the loudest moments? Because volume finds the scare and not the clip. The audio peak tells Vizard Agent where to look, and the frame a beat later — your reaction — is what decides whether the moment is worth cutting at all.
Will my face always be visible? Yes, if you say so. Vizard Agent treats it as a layout rule rather than a preference.
How long a session can it handle? Half an hour was routine on this run, and Vizard Agent reads all of it rather than sampling.
Does it keep sixty frames per second? Yes. Gameplay needs it, so the delivery matches the source rate.
Can I get a vertical version? Yes, with the facecam stacked over the game rather than one cropped away.
Why check the reaction frames separately? Because the moment worth clipping is usually a beat after the moment in the game. Looking only at the gameplay finds the scare; looking at your face finds the clip.
Can Vizard Agent add subtitles? Yes, built from the same transcript it already made in order to find the moments in the first place.
Does Vizard Agent check the finished reel? Yes. It reviews the assembled cut frame by frame and inspects the waveform, which is how it caught a clip splitting a sentence on this run.
Does it work for other genres? Yes. Horror, racing, shooters — what changes is what counts as a moment, not the method.