How to switch the camera to whoever is talking in a recorded panel
Upload the recording and tell Vizard Agent to cut to whoever is speaking. Vizard Agent transcribes the episode to find the speaker turns, matches each voice to the face on screen, builds a switching plan that goes full-screen for long answers and back to the grid for quick exchanges, and smooths the flicker out.
What is the short version?
A recorded panel where everyone sits in a grid for an hour is unwatchable, and the fix is the thing a live director does: cut to whoever is talking. Doing that afterwards means knowing who is speaking and which box on screen they occupy.
- Go to Vizard Agent and upload the full recording.
- Say to go full-screen on the speaker and back to the grid for exchanges.
- Supply the intro music and the artwork for the bookends.
What do you need before you start?
The recording and your bookend assets. Vizard Agent works out the speaker turns and the screen layout itself, so no timeline notes are needed, but it needs the music and artwork if the episode is meant to open and close properly.
- The recording. The whole episode, grid and all.
- The music. An intro track, to fade in and out.
- The artwork. Shown at the start and the end.
- The switching rule. Long talker full-screen, exchanges in grid.
- The platform. It sets the frame and the loudness target.
What do you type into Vizard Agent?
Describe the switching rule rather than the timestamps. Vizard Agent needs the policy — what counts as a long turn, what counts as an exchange — and it derives every cut from the transcript rather than from a list you have to write.
Prompt
Variants worth knowing:
- A minimum hold. So the picture does not flicker between speakers.
- Grid as the default. With full-screen reserved for real answers.
- Levels matched. Music under speech at a stated target.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real fifty-six minute panel episode. The step that makes the switching possible at all is the thirteenth: it worked out which person on screen corresponds to which voice in the audio.
- Checks the uploaded video and music track, then looks at the artwork.
- Samples frames across the podcast to see the screen layout and reviews them.
- Downloads the files and transcribes the full episode to find who talks when.
- Inspects the transcript structure and scans the whole video for layout changes between grid and split-screen.
- Works out the speaker turns and who talks longest.
- Measures the exact grid geometry and views the four-person and two-person layouts at full size.
- Identifies which on-screen person matches each voice.
- Finds the exact moment the layout switches and measures the loudness of the music and the speech three separate ways.
- Builds the cut plan — full-screen for long talkers, grid for quick exchanges — then smooths out the flickery short cuts and cleans the plan up.
Step 9's second half is what separates this from a mechanical rule. Cutting on every speaker change produces a picture that flickers during a rapid exchange, so the plan has to be smoothed after it is built rather than applied as written.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1920x1080, H.264, 25fps, 3374.66 seconds, AAC audio. Fifty-six minutes of finished episode — artwork and music fading in at the top, the same at the end, and a picture that follows the conversation instead of sitting on a static grid.
This is the full episode rather than clips cut from it. That is a different deliverable from the usual podcast job, and it is the one that actually gets published on a channel.
When does this not work well?
Automatic switching depends entirely on being able to tell the speakers apart, both in the audio and on the screen. These are the cases where Vizard Agent has less to work with than the job needs, and it will say so rather than guess.
- Overlapping speech confuses turns. Two people talking at once has no single speaker.
- Similar voices are hard to separate. Speaker identification is not infallible.
- A changing layout complicates the geometry. Boxes that move need re-measuring.
- Poor framing survives zooming. A badly composed box is worse full-screen.
- Very short turns should not cut. Say the minimum hold you want.
How do you fix a result that came back wrong?
Point at the timestamp where it goes wrong. Vizard Agent keeps the transcript with its speaker turns, the measured grid geometry, the voice-to-face mapping and the cut plan, so the switching can be re-tuned without re-analysing the whole episode.
- "It cuts too often around 12 minutes." The minimum hold is raised in that stretch.
- "Wrong person on screen." Re-mapped from the voice identification.
- "The music is too loud under the intro." Re-levelled from the measurements taken.
How does Vizard Agent compare to doing it yourself?
By hand this is an hour of episode and a cut every few seconds, made by watching and clicking. Most people give up and publish the static grid, which is why so many recorded panels look like a video call rather than a programme.
| By hand | Vizard Agent | |
|---|---|---|
| Speaker turns | Watch and mark | Derived from the full transcript |
| Who is where | Obvious to you, tedious to log | Voices matched to on-screen positions |
| Switching | Cut by cut | A plan built, then smoothed |
| Levels | Balance by ear | Loudness measured, three ways |
Common questions
How long an episode can Vizard Agent handle? Fifty-six minutes here, transcribed and switched in full. Longer works the same way.
How does it know who is speaking? Vizard Agent transcribes the episode for speaker turns and matches each voice to a face.
Will it cut on every single change? No. Vizard Agent smooths the plan so rapid exchanges stay in the grid.
Can it add the intro music and artwork? Yes, with fades. Vizard Agent places the artwork at both ends of the episode.
What about the audio balance? Vizard Agent measures the music and speech loudness rather than mixing by ear.
Why smooth the switching plan afterwards? Because a rule that cuts to whoever is talking produces a flickering picture during a fast exchange. The plan has to be built first and then relaxed, which is exactly what a live director does instinctively.
Does this work for video calls? Yes. Vizard Agent treats a recorded call in grid view as exactly the same problem.
Can I get clips from the same pass? Yes. Vizard Agent already holds the transcript and the speaker turns from this pass.
What if two people talk over each other? Vizard Agent holds on the grid rather than picking a winner between them.
Is a full episode worth publishing as well as clips? For a channel, yes. Vizard Agent produces both from the same analysis, and the switched full episode is what gives the clips somewhere to send people back to.
Does it check the finished episode? Yes. It reviews the cut plan and the levels before rendering and delivery.