Vizard Agent

How to remove background noise and other voices from a video

Last updated 2026-08-16 · 5 min read

Upload the recording to Vizard Agent and say what should not be in the audio. Vizard Agent transcribes the whole thing first so it knows who speaks when, analyses the background sounds across the recording, and works from word-level timings to remove the noise and the voices you named while leaving your speaker intact.

What is the short version?

Bad audio ruins a video faster than bad picture, and the two common problems are different jobs. Constant noise — traffic, a fan, room hum — is a filtering problem. A second person talking off camera is a timing problem, because the fix has to know exactly when they spoke.

  1. Go to Vizard Agent and upload the recording.
  2. Say what should be removed: the hum, the voices, a specific sound.
  3. Say who or what must stay untouched.

What do you need before you start?

The original recording and a clear description of what is unwanted. Vizard Agent transcribes everything before it removes anything, which means it can act on "the third person's voice" as a real instruction rather than a description of a feeling — but only if you say whose voice belongs in the video.

What do you type into Vizard Agent?

Name the unwanted sound and the wanted one. Vizard Agent can hear the difference between a voice and a hum, and it cannot know which of two people in a room is the subject of your video — that distinction is the whole instruction, and it takes one clause.

Prompt

Remove all the unwanted background noise and extra sounds from this, including [the third person's voice]. Keep [the main speaker] clean, and improve the quality if you can.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real cleanup. The whole first half is listening and locating: it builds a picture of what is in the audio and precisely when, before touching any of it.

  1. Analyses the video's properties and structure.
  2. Transcribes the audio to see what the dialogue actually is.
  3. Analyses the audio and the voices to locate the background noise and anything else present.
  4. Analyses the end of the recording specifically for unwanted voices or sounds.
  5. Performs a detailed pass over all the background voices and noise across the full recording.
  6. Reads the word-level transcript for precise timings.
  7. Extracts a grid of frames across the video to inspect the picture as well.
  8. Inspects that grid for the speaker's face, lighting and movement.

Step 6 is what makes voice removal possible rather than approximate. Word-level timings turn "the other person talking" into a set of exact intervals, and an interval can be treated where an impression cannot.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 5.02 seconds, AAC audio. The picture geometry was preserved, since cleaning the audio is not a reason to change the frame, and the delivered file carries the cleaned track in place of the original.

Ask for the cleaned audio on its own and Vizard Agent will hand you the track, which is useful when the edit is happening somewhere else.

It is not instant, and this category was not separately measured. Comparable work runs a median of 28 to 38 minutes end to end. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.

When does this not work well?

Separating sounds that were recorded together has a ceiling. Vizard Agent can lift a voice out of noise and cut a voice out of a gap, and it cannot unmix two people speaking over the top of each other.

How do you fix a result that came back wrong?

Point at the timestamp. Vizard Agent keeps the transcript, the word-level timings and its analysis of what is in the audio, so a second pass targets a specific interval rather than re-analysing the recording, and the parts you were happy with stay untouched.

How does Vizard Agent compare to doing it yourself?

By hand this is a noise-reduction plugin and a lot of scrubbing: find every place the other person speaks, cut or duck each one, then adjust the reduction until the voice stops sounding metallic. The finding is the slow part, and it is done by ear in real time.

By hand Vizard Agent
Finding the unwanted voice Listen through, mark by ear Transcribed with word-level timings
Noise reduction Set a level and hope Analysed across the whole recording
Checking the result Listen again, end to end Analysed and reported
A second pass Start from the original Targets the interval you name

Common questions

Can it remove one person and keep another? Yes, when they do not talk at the same time. Overlapping speech is where this stops working cleanly.

Will my voice sound processed? Not unless the noise is heavy enough to require it. Say "keep it natural" and Vizard Agent stays conservative.

Can I get just the audio file? Yes. Ask for the cleaned track on its own if the edit is happening elsewhere.

How long a recording can it handle? Long recordings are fine. Vizard Agent transcribes and analyses the whole file before treating any of it, so the work scales with duration rather than failing at a threshold.

Can it fix the picture at the same time? Yes, and doing both in one job is cheaper than two, because Vizard Agent inspects the footage once.