How to remove background noise and other voices from a video
Upload the recording to Vizard Agent and say what should not be in the audio. Vizard Agent transcribes the whole thing first so it knows who speaks when, analyses the background sounds across the recording, and works from word-level timings to remove the noise and the voices you named while leaving your speaker intact.
What is the short version?
Bad audio ruins a video faster than bad picture, and the two common problems are different jobs. Constant noise — traffic, a fan, room hum — is a filtering problem. A second person talking off camera is a timing problem, because the fix has to know exactly when they spoke.
- Go to Vizard Agent and upload the recording.
- Say what should be removed: the hum, the voices, a specific sound.
- Say who or what must stay untouched.
What do you need before you start?
The original recording and a clear description of what is unwanted. Vizard Agent transcribes everything before it removes anything, which means it can act on "the third person's voice" as a real instruction rather than a description of a feeling — but only if you say whose voice belongs in the video.
- The original audio. Anything already noise-reduced has been through one pass, and stacking two makes voices sound underwater.
- What should go. Traffic, a fan, a dog, a person talking off camera. The more specific, the better the result.
- Who must stay. Name the speaker who belongs in the video, especially when more than one person is audible.
- Roughly when, if you know. "The other voice is near the end" saves Vizard Agent scanning for it.
- Whether picture matters too. Recordings with bad audio often have bad light as well, and both can be handled in one job.
What do you type into Vizard Agent?
Name the unwanted sound and the wanted one. Vizard Agent can hear the difference between a voice and a hum, and it cannot know which of two people in a room is the subject of your video — that distinction is the whole instruction, and it takes one clause.
Prompt
Variants worth knowing:
- Noise only. "Just the room hum, leave everything else." Safer on recordings with more than one wanted voice.
- A specific moment. "There is a door slam around 0:40." Vizard Agent will target it rather than treating the whole file.
- Audio and picture together. Ask for both and the transcription work is shared between the two jobs rather than done twice.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real cleanup. The whole first half is listening and locating: it builds a picture of what is in the audio and precisely when, before touching any of it.
- Analyses the video's properties and structure.
- Transcribes the audio to see what the dialogue actually is.
- Analyses the audio and the voices to locate the background noise and anything else present.
- Analyses the end of the recording specifically for unwanted voices or sounds.
- Performs a detailed pass over all the background voices and noise across the full recording.
- Reads the word-level transcript for precise timings.
- Extracts a grid of frames across the video to inspect the picture as well.
- Inspects that grid for the speaker's face, lighting and movement.
Step 6 is what makes voice removal possible rather than approximate. Word-level timings turn "the other person talking" into a set of exact intervals, and an interval can be treated where an impression cannot.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 5.02 seconds, AAC audio. The picture geometry was preserved, since cleaning the audio is not a reason to change the frame, and the delivered file carries the cleaned track in place of the original.
Ask for the cleaned audio on its own and Vizard Agent will hand you the track, which is useful when the edit is happening somewhere else.
It is not instant, and this category was not separately measured. Comparable work runs a median of 28 to 38 minutes end to end. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.
When does this not work well?
Separating sounds that were recorded together has a ceiling. Vizard Agent can lift a voice out of noise and cut a voice out of a gap, and it cannot unmix two people speaking over the top of each other.
- Overlapping speech is the hard limit. If the unwanted voice talks while your speaker talks, removing one damages the other. The honest fix is to cut that passage.
- Heavy noise reduction makes voices sound processed. There is a point where the cure is more noticeable than the disease, and Vizard Agent will stop short of it unless you push.
- Wind and clipping cannot be undone. Wind on a microphone destroys the signal rather than covering it, and clipped audio has no waveform left to recover.
- Room echo is not noise. Reverb is your voice arriving twice. It can be reduced a little and not removed.
- A cleaned track can still sound thin. Removing everything unwanted sometimes leaves a voice with nothing around it, which reads as sterile rather than clean. A small amount of room tone put back is often the better finish.
- Consent still applies. Removing someone's voice from a recording does not settle whether they agreed to be recorded.
How do you fix a result that came back wrong?
Point at the timestamp. Vizard Agent keeps the transcript, the word-level timings and its analysis of what is in the audio, so a second pass targets a specific interval rather than re-analysing the recording, and the parts you were happy with stay untouched.
- "There is still a voice at 0:38." Treated as an interval, using timings it already has.
- "My speaker sounds processed now." Vizard Agent backs off the reduction.
- "You removed a sound I wanted." Restored from the original track.
How does Vizard Agent compare to doing it yourself?
By hand this is a noise-reduction plugin and a lot of scrubbing: find every place the other person speaks, cut or duck each one, then adjust the reduction until the voice stops sounding metallic. The finding is the slow part, and it is done by ear in real time.
| By hand | Vizard Agent | |
|---|---|---|
| Finding the unwanted voice | Listen through, mark by ear | Transcribed with word-level timings |
| Noise reduction | Set a level and hope | Analysed across the whole recording |
| Checking the result | Listen again, end to end | Analysed and reported |
| A second pass | Start from the original | Targets the interval you name |
Common questions
Can it remove one person and keep another? Yes, when they do not talk at the same time. Overlapping speech is where this stops working cleanly.
Will my voice sound processed? Not unless the noise is heavy enough to require it. Say "keep it natural" and Vizard Agent stays conservative.
Can I get just the audio file? Yes. Ask for the cleaned track on its own if the edit is happening elsewhere.
How long a recording can it handle? Long recordings are fine. Vizard Agent transcribes and analyses the whole file before treating any of it, so the work scales with duration rather than failing at a threshold.
Can it fix the picture at the same time? Yes, and doing both in one job is cheaper than two, because Vizard Agent inspects the footage once.