How to use a separate audio track for each podcast guest
Send the camera files and one audio track per guest. Vizard Agent syncs each track to the picture, verifies the alignment holds to the end of the episode, works out who is on which channel, cuts between cameras by who is actually speaking, and gates the bleed out of each microphone.
What is the short version?
A separate track per person is not just cleaner audio. It is a record of who was talking at every instant of the episode — which is exactly the information a multicam edit needs and normally has to guess at from the picture.
- Go to Vizard Agent with the cameras and every audio track.
- Say there is one track per guest plus a combined one.
- Ask it to cut between cameras based on the tracks.
What do you need before you start?
The camera files and the individual per-person tracks. A combined track is useful to Vizard Agent as a reference, but the separate ones are what earn their keep — without them, "who is speaking" is a guess made from watching mouths move.
- The camera files. However many angles.
- One track per person. The whole point.
- The combined track. Handy as a fallback.
- Who is who. Or let it identify them.
- Your logo and format. Corner, resolution, platform.
What do you type into Vizard Agent?
Say what the tracks are. A folder of files does not announce that channel three is the guest and channel four is both of them together — and that mapping is the thing that makes the rest of the edit automatic.
Prompt
Variants worth knowing:
- "One channel per guest." The mapping, stated.
- "Cut on who is speaking." What the tracks are for.
- "Clean the bleed." The other thing they enable.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real two-camera, three-channel interview episode delivered for YouTube. Steps three to six are the point where the separate tracks stop being a filing convention and start doing real editorial work for Vizard Agent.
- Checks the project folder and looks at the supplied logo.
- Measures the audio channels and checks the file sizes.
- Transcribes the channels and syncs the first camera to the audio.
- Measures the drift between camera and audio precisely.
- Searches the transcripts for microphone notes and long pauses.
- Checks whether anyone spoke during the pauses the transcript missed.
- Identifies who is on each channel and looks at frames from the camera.
- Confirms the guest's name from the transcript.
- Syncs the audio channels to the camera timeline properly.
- Verifies the sync accuracy across the whole episode, not just the start.
- Builds the cut list between cameras from who is speaking.
- Adds reaction shots of the listener during long stretches.
- Fixes the cut around the microphone note and at the episode's end.
- Checks the shots either side of that cut and hides it in a camera change.
- Builds the episode's soundtrack from the microphones.
- Places the logo and checks it on a real frame.
- Gates the bleed out of each microphone, keyed to who is speaking.
- Re-mixes with the cleaned channels and prepares a test passage.
- Listens to that passage for gating artifacts.
- Uploads the cleaned microphone so you can hear the difference.
Step eleven is the payoff. With per-guest tracks the cut list is derived rather than judged — the camera changes when the speaker changes, because the tracks say when that happened, to the frame.
Step seventeen is the other half, and it needs the same information. A gate that opens on volume alone chops the beginnings of words; one keyed to which channel actually has the speaker on it can be much more aggressive without damaging anything.
What does the result look like?
An episode that cuts to whoever is talking at that moment, with reaction shots wherever somebody listens for a long stretch, and audio in which each microphone carries its own person and very little of anybody else in the room with them.
The cutting feels attentive rather than mechanical. That is a side effect of the cut list coming from the conversation rather than from a rhythm imposed on top of it.
When does this not work well?
Everything here depends on the tracks and the picture being alignable at all, and on each track genuinely belonging to exactly one person. Where either of those breaks down, per-track editing offers Vizard Agent nothing over an ordinary multicam edit.
- Tracks that drift. Long recordings from separate devices creep apart.
- One shared microphone. No per-person information to use.
- Heavy crosstalk. Both channels have both people on them.
- A missing camera file. Nothing to cut to when that person speaks.
- Very aggressive gating. It can clip quiet agreement sounds.
How do you fix a result that came back wrong?
Say whether the problem is in the cutting or in the sound. Vizard Agent keeps the sync measurements, the speaker mapping and the cut list as separate things, so a mis-timed camera change and an over-gated microphone end up being two entirely different repairs.
- "It cuts too often." Minimum shot length raised.
- "You missed a reaction." Listener shots added in that stretch.
- "The gate is chopping her." Loosened on that channel only.
- "It drifts near the end." Re-verified and re-synced across the episode.
How does Vizard Agent compare to doing it yourself?
Multicam editors do exactly this, and doing it by hand is a known, respectable workflow. What it costs is attention: syncing, mapping, building the cut list, then gating each channel — for every episode, every week, before any creative decision gets made.
| By hand | Vizard Agent | |
|---|---|---|
| Sync | Aligned at the start | Verified across the whole episode |
| Who is speaking | Read from the picture | Read from the tracks |
| The cut list | Built by hand | Derived from the channels |
| Bleed | Gated on volume | Gated on who is actually speaking |
| Reaction shots | Remembered | Added on long stretches |
Common questions
How many tracks can it take? As many as you recorded. Vizard Agent maps each to a person.
Do I need the combined track? No, but it is a useful fallback. Send it if you have it.
Will it know who is who? Vizard Agent works it out from the transcript and the picture, then confirms names.
What if the tracks drift? Vizard Agent measures the drift and corrects it across the episode.
Can it cut on something other than speech? Yes. Say if you want reaction-led or rhythm-led cutting instead.
What is bleed? Each microphone picking up the other person. Vizard Agent gates it out.
Will gating sound unnatural? Vizard Agent tests a passage and listens for artifacts before applying it.
Can I hear the difference? Yes. Ask Vizard Agent to upload the cleaned channel on its own.
Does it add reaction shots? Yes. Vizard Agent adds them during long single-speaker stretches.
What about the logo? Vizard Agent places it where you say and checks it on a real frame.
Can it deliver 4K? Yes. Say the platform and Vizard Agent works to it.
What if one camera is missing? Vizard Agent tells you and cuts from what exists.
Does it work for three or more guests? Yes, and the benefit grows with each person.
Why not just cut on the picture? Because mouths move for reasons other than talking, and a track knows the difference.