Vizard Agent

How to cut around the moments a guest's microphone dropped out

Last updated 2026-09-05 · 8 min read

Tell Vizard Agent that a guest's microphone cut out during the recording. It searches the transcript for the moments you mention, scans the rest of the show for silent gaps that look like the same fault, reads word-level timings through each choppy stretch, and builds the cut plan so every removal lands between words.

What is the short version?

A dropped microphone does not sound like silence when you are in the room, so hosts usually catch two or three of them and miss the rest. The transcript catches all of them, because a dropout leaves a hole in the words rather than a pause between them.

  1. Go to Vizard Agent with the recording.
  2. Say whose microphone cut out, and roughly when if you know.
  3. Ask it to find every other dropout as well.

What do you need before you start?

Very little beyond the recording itself. A rough timestamp helps as a starting point but is not required, and it is genuinely better to say "a few times during the interview" than to guess at times you are not sure about and have the search anchored to the wrong place.

What do you type into Vizard Agent?

Ask the second question as well as the first. "Clean up these three moments" gets you three moments cleaned; "find any others" is what turns a patch into a pass, and on a long interview it almost always finds something you had not heard.

Prompt

A few times during the interview my guest's mic cut off. Can you clean those parts up? Please also scan the whole recording for any other dropouts I did not notice, and cut between words so her answers still make sense.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real podcast episode where a guest's microphone failed intermittently during the interview. The first pass follows what the host actually reported hearing; the second pass, the sweep across everything else, is the one that matters.

  1. Searches the transcript for the moments the host described.
  2. Reads around each one to see how much speech is affected.
  3. Scans the whole show for silent gaps that look like the same fault.
  4. Downloads the source locally to examine a suspicious stretch closely.
  5. Locates the word-level transcript rather than the sentence one.
  6. Reads word timings through the choppy stretch to see where words are missing.
  7. Re-reads that stretch to separate lost words from ordinary pauses.
  8. Builds the cut plan from the word timings, dropouts removed.
  9. Test-renders a section and checks the result on frames and waveform.
  10. Has the cut sections watched back for speaker accuracy and smoothness.
  11. Rebuilds the episode around the shortened answers.
  12. Re-checks the open and the close before uploading.

Step three is what separates this from a patch job. A host hears the dropouts that interrupted them mid-question and misses the ones during their own talking, so the sweep looks at the whole recording for the same signature rather than only at the reported times.

Step six is the discrimination that makes the cuts safe. A pause has room tone in it and a dropout does not, and at word level a dropout shows up as words that stop and restart without the gap that a person leaves when they are thinking.

What does the result look like?

An episode where the damaged seconds are gone and the answers still read as answers. The guest's replies are slightly shorter than they were recorded, the joins land between words, and nothing in the conversation stops making sense.

The usual reaction is that nobody notices anything was wrong.

When does this not work well?

Cutting only works when there is something on either side worth keeping. If the microphone was down for most of an answer, or the answer's key sentence is the part that vanished, no amount of clean cutting brings it back.

How do you fix a result that came back wrong?

Say what you hear and where. Vizard Agent keeps the word-level transcript and the cut plan built from it, so a note like "that one is still there" or "you took too much" changes a boundary in the plan rather than starting the episode again.

How does Vizard Agent compare to doing it yourself?

By hand this is done by listening to the whole episode with a finger on the spacebar, which is slow and finds only what you can hear on the day. The cuts then get placed at whatever looked like a gap on the waveform, which is often mid-word.

By hand Vizard Agent
Finding dropouts The ones you remember The whole recording swept
Telling them from pauses By ear Room tone and word timings
Where the cut lands Nearest visible gap Between words
Whether the answer survives Hoped Checked against the transcript
Finding a second fault Usually not Same signature, same sweep

Common questions

Do I need to know when it happened? No. Vizard Agent scans the whole recording either way.

Will the guest sound choppy? Not if the cuts land between words, which is what Vizard Agent plans for.

Can it fix the audio instead of cutting? Only where something is there to repair. Vizard Agent will say which.

What if there were separate tracks? Better. Vizard Agent can cut one track and leave the others running.

Does the video get cut too? Yes, in step with the audio, so Vizard Agent preserves lip sync.

Will it tell me how many it found? Yes. Ask Vizard Agent for the list of dropouts and their times.

What about the host's microphone? Same sweep. Vizard Agent checks every speaker, not just the one you named.

Can it keep the question and drop the answer? Yes. Tell Vizard Agent which side matters.

Does this change the episode length? Slightly shorter. Vizard Agent reports the new runtime.

What if a dropout is mid-sentence? Vizard Agent cuts to the nearest clean word boundary either side.

Can it work on a video call recording? Yes. That is where this fault is most common.

Will captions still match? Yes. Vizard Agent regenerates them from the cut audio.

Can I hear the removed sections? Yes. Ask Vizard Agent to export them separately.

Why look at words rather than the waveform? Because a dropout and a thoughtful pause look almost identical in a waveform and completely different in a word-timed transcript.