Vizard Agent

How to get a clean speaker crop out of a streaming platform layout

Last updated 2026-09-04 · 7 min read

Tell Vizard Agent which person's box you want out of the recording. It finds the webcam tiles by their border colour, crops one, looks at the result to check whether the platform's frame is still visible, and then crops inside the border until what is left is only the speaker.

What is the short version?

A recording from a streaming studio is not footage of a person. It is a picture of a layout — boxes, borders, name bars, a logo — and cropping to the box still leaves you with the box's decoration in the shot.

  1. Go to Vizard Agent with the recording.
  2. Say whose webcam tile you want, and in what frame.
  3. Ask it to check the crop actually has no border in it.

What do you need before you start?

The recording as it came out of the platform. Nothing needs preparing — Vizard Agent works out the layout by looking at it, which is more reliable than being told coordinates that turn out to be from a different scene.

What do you type into Vizard Agent?

Say which person, not which coordinates. Layouts move — a guest joining or a screen share starting will reshuffle the boxes — so an instruction phrased around a person survives the recording changing shape and one phrased around a rectangle does not.

Prompt

Pull the host's webcam feed out of this recording as a clean vertical clip. Make sure the crop doesn't include the platform's border or name bar — check it on the actual frames.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real podcast recorded in a streaming studio with gold-bordered speaker tiles. Steps four to eight are the loop that actually matters here — crop it, look at what came out, then crop tighter until the platform is gone.

  1. Inspects the recording for its duration and stream details.
  2. Extracts frames across the video to see the layout and who is where.
  3. Extracts a full-size frame to get the exact coordinates of the host's box.
  4. Detects the bordered speaker boxes by segmenting on the border colour.
  5. Crops the host's webcam area and checks whether it is clean.
  6. Looks at the cropped image to see if the platform frame survived.
  7. Crops inside the border to get the feed on its own.
  8. Looks again to confirm no border remains.
  9. Checks segments where the content changes to see what is on screen.
  10. Concatenates the pieces losslessly and verifies the streams.
  11. Confirms the audio survived the operation.
  12. Renders the clips in the layout you asked for.
  13. Extracts timeline frames to verify subtitles, layout and overlays.

Steps five to eight are the whole method, and the reason it is a loop rather than a calculation. A crop derived from measured coordinates is nearly always a few pixels generous, so a sliver of the platform's border survives at one edge — invisible in a thumbnail and obvious at full screen.

Step four is why this works on layouts nobody described. The tiles are found by their own border colour rather than by a template, so a studio you have never used before is handled the same way as a familiar one.

What does the result look like?

A clean shot of one person, filling the frame, with no border, no name bar and no trace of the platform it was recorded in. It reads as camera footage rather than as a cropped screenshot of a call.

Where you asked for both speakers, you get two independent sources from one recording — which is what makes it possible to cut between them afterwards rather than living with the layout the studio chose.

When does this not work well?

The method assumes each person occupies their own rectangle in the frame. Layouts that overlap, blur or animate the tiles never give Vizard Agent a clean box to take, and some recordings arrive already composited past the point where anything can be separated out of them.

How do you fix a result that came back wrong?

Say what is still in the frame. Vizard Agent keeps the detected box coordinates and the crop results, so tightening one edge or switching to the other speaker is an adjustment rather than a fresh analysis of the layout.

How does Vizard Agent compare to doing it yourself?

By hand this is measuring a rectangle in a still and applying it to the whole recording, which works until the layout changes or your rectangle was two pixels too wide. Both of those are found late, usually by looking at the published clip.

By hand Vizard Agent
Finding the box Measured on one frame Detected by border colour
Checking the crop Trust the numbers Looked at, then tightened
Layout changes Break the crop Handled per section
Two speakers Two manual passes Both from one analysis
Audio Easy to lose Verified after the operation

Common questions

Does it need to know the platform? No. Vizard Agent finds the boxes by looking rather than by template.

Can it do both speakers? Yes. Vizard Agent produces each as its own clean source.

What if the layout changes? Vizard Agent crops each section separately rather than forcing one rectangle.

Will the audio come with it? Yes, and Vizard Agent verifies it survived the crop.

Can it make it vertical? Yes. Vizard Agent reframes as it crops; that is usually the point.

What about the name bar? Cropped out. Vizard Agent checks the result rather than assuming.

Is the quality good enough? It depends on the tile size. Vizard Agent tells you when it is too small.

Can it upscale the crop? Vizard Agent can, though a small tile stays a small tile underneath.

What if the borders are the same colour as the background? Harder. Vizard Agent falls back to edge detection and shows you the result.

Can it keep the original layout too? Yes. Ask and Vizard Agent delivers both.

Does it work for a screen share section? Differently — nobody is in a box. Say what you want from those parts.

Can it cut between the two speakers? Yes. Once Vizard Agent has each as a clean source, it can cut freely.

How does it know who is the host? Tell Vizard Agent, or it will show you the boxes and ask.

Why not just crop to the coordinates? Because a measured rectangle is generous by a few pixels, and those pixels are the platform's border at full screen.