Vizard Agent

How to put filler gameplay under a talking clip to hold attention

Last updated 2026-09-22 · 8 min read

Tell Vizard Agent that the gameplay in the lower half is a background rather than a subject. The talking is the content — the clips are chosen from the transcript, the camera is cropped clear of the stream's own overlay, and the filler is only there to occupy the eye while someone listens.

What is the short version?

You want the format where a face talks in the top half and something endlessly watchable runs underneath it. Getting it wrong means the bottom half starts competing with the story rather than supporting it. Vizard Agent treats the filler as furniture and the talking as the content.

  1. Tell Vizard Agent the layout and that the filler is background.
  2. Let it pick the moments from what was said.
  3. Check the camera crop excludes the chat and the overlay.

What do you need before you start?

The stream recording and the filler footage, plus any title cards you want at the head and tail. If the filler is not in the recording, supply it — the point of this format is that the bottom half is not from your stream at all.

Say what kind of talking moment you want. In the session behind this article the brief listed them explicitly: unexpected stories, funny situations, hot takes, arguments, admissions, sudden turns. That list is what the transcript gets searched against, and it is much more useful than "the best bits".

What do you type into Vizard Agent?

Describe the two halves and say which one matters. The second part is the one people leave out, and without it the filler gets treated as footage to cut to rather than as a texture to run underneath.

Vertical, streamer's camera on top, parkour gameplay underneath. The talking is the content — the gameplay is only there to hold the eye and must not pull attention off the speech.

That clause about not pulling attention is the whole design brief for the lower half. It rules out cutting the filler on the beat, putting anything dramatic in it, or letting its own sound through.

What does Vizard Agent actually do?

It works from the transcript and treats the picture as assembly. The stream is transcribed, sorted by topic, and searched for the hooks you described; only once the moments are chosen does the layout get built around them.

In that session the sequence ran:

Then a long run on the captions, which is where this format spends its time: the engine kept breaking words across lines, and the fix was eventually to build whole phrases with exact timings so the automatic wrapping never ran at all.

What does the result look like?

Clips where you watch the face and read the captions, and the bottom half is simply there. The speaker is framed with no chat window or overlay in shot, the filler runs continuously rather than cutting, and each clip stands alone with its own opening and closing card.

The captions carry more weight here than usual, because a lot of this format is watched with the sound off in the first second. Whole words, no mid-word breaks, no overlapping lines — those are the things that get checked on the delivered file rather than in the timeline.

When does this not work well?

When the talk needs the picture to make sense. A story about something that happened on screen cannot be told over unrelated parkour, and this format quietly removes your ability to show anything at all. Those moments want the normal treatment instead.

It also does not rescue weak material. The filler buys a second or two of attention, not a minute; if the story is not interesting the clip still ends early. Vizard Agent will say when the candidates are thin rather than making five clips out of three good ones.

And a stream whose camera window is small, or partly covered by overlays and alerts, leaves nothing clean to crop out of it. Enlarging a small webcam box to fill half a vertical frame looks exactly as soft as it sounds, so that one is worth fixing in your streaming layout rather than in the edit afterwards.

How do you fix a result that came back wrong?

If the filler is distracting, say so. Vizard Agent can dim it, slow it, or pick a less eventful section — the fix is almost never to make it smaller, because the layout proportions are what make the format read.

If the camera crop includes chat or an overlay, name the edge. The crop is set from a measured frame, so a specific edge is a specific adjustment.

If the captions break words, ask for whole phrases with fixed timings rather than a different size. Resizing pushes the problem around; removing the automatic wrapping ends it, which is what that session eventually did.

How does Vizard Agent compare to doing it yourself?

The format is easy to assemble and hard to select for. Anyone can stack two videos; the work is in finding the forty seconds of a three-hour stream that hold up with no visual support at all, which is a reading job rather than a watching one.

That is where working from the transcript pays. Sorting the conversation by topic and hunting hooks across the whole length finds the moments evenly, rather than favouring whatever happened while you were still paying attention.

Common questions

Where does the filler come from? You supply it, or Vizard Agent sources something neutral. It is not from your stream.

Should the filler have sound? No. Vizard Agent mutes it so only the speech is heard.

What split works best? Roughly half and half, with the camera on top. Vizard Agent measures real heights rather than fractions.

Can I use my own title cards? Yes. Send them and Vizard Agent places them at head and tail.

How long should each clip be? As long as the story needs. Vizard Agent will not pad to a target.

Will the chat be visible? Not if the crop is set properly. Vizard Agent verifies on a frame.

How does it choose the moments? Vizard Agent reads the transcript and matches it against the kinds of moment you described.

Can it make several at once? Yes. Vizard Agent builds them as a set with consistent styling.

What about the captions? Whole words, no overlaps. Vizard Agent checks that on the delivered file.

Can the filler change between clips? Yes, if you supply more than one, though Vizard Agent will suggest keeping it consistent.

Does the speaker need to be centred? In their half, yes. Vizard Agent crops to keep the face framed.

What resolution should I ask for? Whatever the platform needs. Vizard Agent will say what the source supports.

Can it add a hook card at the start? Yes, that is the opening title card, and Vizard Agent checks the wording fits.

Will the filler loop visibly? Not if there is enough of it. Vizard Agent picks one continuous stretch rather than repeating a short one.

Can I put the gameplay on top instead? Yes, though the face reads better in the upper half. Vizard Agent will build it either way.