Vizard Agent

How to stack two interview speakers in one vertical frame

Last updated 2026-09-04 · 7 min read

Ask Vizard Agent for a double frame and it builds one vertical shot with both speakers stacked. It measures where each person sits in the wide recording, draws a coordinate grid over a real frame to pin the crop boxes exactly, and re-measures for every scene where they move.

What is the short version?

A two-person interview filmed as a wide shot does not crop down to vertical. Centre the frame and you get the gap between them; crop to one person and you lose the other entirely. Stacking gives you both, each one filling their own half of the frame.

  1. Go to Vizard Agent with the wide interview footage.
  2. Ask for the clips in a double frame — one speaker above the other.
  3. Say where the location changes, if it does.

What do you need before you start?

The wide recording and a note about the settings. A conversation that moves from a car to a street is really two framing problems, and saying so up front saves a round where the second half is cropped to where people were sitting in the first.

What do you type into Vizard Agent?

Ask for a double frame by name. It is the term the format goes by, and it is unambiguous in a way that "put them both in shot" is not — that could reasonably mean a wide clip letterboxed into a vertical frame, which is the thing you are trying to avoid.

Prompt

Cut my interview into the double frame — three of the best short clips for TikTok, 30 to 60 seconds each. Host on top, guest underneath, and re-measure the crops for the outdoor scenes.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real ten-minute interview cut into three vertical clips. Steps six to nine are the measuring, and they are why the faces end up centred in their halves rather than approximately in them.

  1. Checks the footage's format and transcribes the audio.
  2. Views a sample frame to see the layout and who is where.
  3. Builds a contact sheet across the whole interview.
  4. Reviews the speaker positions as they change through the recording.
  5. Reads the full transcript and picks the candidate clips.
  6. Extracts frames from the first scene to determine the crop coordinates.
  7. Measures the speaker positions on a full-resolution frame.
  8. Calculates the double-frame crop boxes.
  9. Generates a layout preview and checks the framing and balance.
  10. Refines the crop coordinates for both people.
  11. Draws a pixel coordinate grid over the frame to find exact locations.
  12. Extracts frames from the outdoor scenes where the layout differs.
  13. Views the street frame to see the multi-person layout there.
  14. Tests the precise split and reviews it.
  15. Checks the timestamps around the transitions between scenes.
  16. Generates word-timed transcripts for each clip.
  17. Builds styled captions and tests the burn on the composition.
  18. Renders a test slice with captions and a headline, and reviews it.

Step eleven is the trick worth stealing. Rather than estimating coordinates from a description of the frame, it draws a numbered grid over an actual frame and reads the positions off it — which turns "somewhere on the left, about a third down" into two numbers.

Step twelve is the reason this is not a single measurement. People sit differently in a car than they stand in a street, and a crop that was perfect for the first scene puts someone's forehead at the edge of the frame in the second.

What does the result look like?

A vertical clip with both speakers stacked, each properly centred in their own half, and captions sitting at the join where they cover neither face. When the conversation moves to a new location, Vizard Agent moves the framing with it rather than reusing the first crop.

Nobody is cut off and there is no dead space between them, which is what happens when a wide two-shot is squeezed into a vertical frame instead of being rebuilt as two.

When does this not work well?

Stacking halves the height available to each person, so Vizard Agent needs a recording with enough resolution to survive that crop, and a layout in which each person can actually be isolated from the other without the two boxes overlapping.

How do you fix a result that came back wrong?

Say which half of the frame and which scene you mean. Vizard Agent keeps the measured crop coordinates separately for every scene it framed, so re-framing one person in one location leaves all the other crops in the clip exactly as they were delivered.

How does Vizard Agent compare to doing it yourself?

By hand this is two crop boxes per scene, typed as numbers, checked by rendering and looking. It is not difficult and it is exactly the kind of work where one scene gets measured properly and the other two get eyeballed.

By hand Vizard Agent
Finding the coordinates Estimated from the frame Read off a drawn pixel grid
A second location Often reuses the first crop Re-measured for each scene
Checking the balance Render and look Layout preview before rendering
Captions Placed and hoped Tested on the composition
Three clips Three manual passes One measurement set

Common questions

What is a double frame? Two speakers stacked in one vertical frame, each filling half. Vizard Agent builds it from a wide shot.

Do I need two cameras? No. Vizard Agent crops both halves out of one wide recording.

What if the location changes? Vizard Agent measures each scene separately rather than reusing one crop.

Can I choose who goes on top? Yes. Name them and Vizard Agent fixes the order.

Will the captions cover a face? No. Vizard Agent places them at the join and tests the result.

What about a third person? There is no room. Vizard Agent will suggest cutting between them instead.

Does the resolution hold up? It depends on the source. Vizard Agent tells you when a half is too soft.

Can it follow people who move? Within limits. Tell Vizard Agent if there is a lot of movement.

Can it cut between single shots instead? Yes, and sometimes that is better. Ask Vizard Agent for both.

Does it pick the clips too? Yes. Vizard Agent chooses from the transcript unless you name the moments.

Can it do this for a podcast layout? Yes, though a bordered studio layout is a different crop problem.

Will it work on a webm recording? Yes. Vizard Agent probes the format first.

Can I get a wide version too? Yes. Vizard Agent renders one from the same clips.

Why not just letterbox the wide shot? Because the two faces end up small and centred with dead space around them, which is exactly what vertical formats punish.