Vizard Agent

How to keep only the sounds that matter from your own footage

Last updated 2026-09-10 · 8 min read

Tell Vizard Agent which sounds in your footage are worth keeping and what should happen everywhere else. It listens through the source for those moments, measures your voiceover against the natural sound underneath it, and lowers or drops the original track wherever the thing you named is not happening.

What is the short version?

Muting your footage loses the sounds that made it worth filming. Keeping all of it puts room noise under your narration for the whole video. What you actually want is a short list of moments kept and everything else out of the way.

  1. Go to Vizard Agent with the footage and your voiceover.
  2. Name the sounds that must survive.
  3. Say what the rest of the track should do.

What do you need before you start?

The footage with its original audio and a description of the sounds by what causes them. Naming the action — a box opening, a lid coming off, a tool starting — is what lets those moments be found; naming a quality, like "the satisfying bits", is not something a track can be searched for.

What do you type into Vizard Agent?

Name the action rather than the feeling it gives you. "Keep the unboxing sounds" identifies a set of moments that exist at particular times in the file; "keep the good bits of audio" is a judgement nobody, including Vizard Agent, can apply consistently across a long clip.

Prompt

Do not mute the original sound — I want the boxes opening and the packaging. Lower the original track in the empty stretches where there is no unboxing happening, and keep my voiceover exactly as it is. Actually, take out everything except the unboxing sounds and my voiceover.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real unboxing reel that had to keep its own sounds underneath a narration the client supplied. Notice how much listening to the source happens before a single level gets set anywhere.

  1. Inspects the uploaded footage and maps the shots across it.
  2. Checks the source audio waveform end to end.
  3. Measures the source speech and the room audio separately.
  4. Listens through the footage for usable speech and audio texture.
  5. Finds the moments where the named action actually happens.
  6. Checks the supplied voiceover and transcribes it.
  7. Measures the voiceover recording before mixing anything.
  8. Compares the voiceover's level against the natural sound underneath.
  9. Builds the mix with the kept moments and the rest lowered.
  10. Renders versions with and without the narration.
  11. Checks the waveform of each delivered version.
  12. Confirms the final levels on both copies.

Step four is the step that decides what is possible. A track can be searched for the moments you named, and it can also be listened to for texture worth keeping that you did not think to ask for — a zip, a lid, a page turning — which is usually where the good material is.

Step eight is what keeps it listenable. A voiceover set to a fixed level will be buried by a loud packaging sound and will tower over a quiet one, so the two get measured against each other rather than each against a target.

What does the result look like?

A track where the sounds you filmed for are present and everything between them has stepped back, with your narration sitting at a level that works over both. You get it with the voiceover and without it, from the same build.

The room noise you never noticed while filming is simply not there any more.

When does this not work well?

The sounds have to be separable in time. If the action you want happens under continuous talking, or the room noise is as loud as the thing you are keeping, then lowering the track around it removes the thing as well.

How do you fix a result that came back wrong?

Say which sound is missing or which stretch is still noisy. Vizard Agent keeps the map of the moments it found and the measurements between the layers, so a correction adjusts one region rather than the whole mix.

How does Vizard Agent compare to doing it yourself?

By hand the choice is usually made once for the whole clip — mute it, or leave it — because doing it per moment means finding every moment. The mix then gets set against a target level rather than against what is actually underneath the voice at that point.

By hand Vizard Agent
The decision Mute or keep, once Per moment, by what happens
Finding the moments Scrubbing The action searched for
Extra texture Missed Listened for and offered
Voiceover level Set to a target Measured against the sound beneath it
Versions One With and without narration

Common questions

Can it keep only certain sounds? Yes. Name them by the action and Vizard Agent finds them.

Will my voiceover be touched? No, not if you say so. Vizard Agent only sets its level.

What about the stretches between? Lowered or silent, whichever you ask Vizard Agent for.

Can it find sounds I did not mention? Yes. Vizard Agent listens for usable texture and will tell you.

Do I get a version without narration? Yes. Vizard Agent builds both from the same assembly.

Can it add sounds that are missing? Yes. Vizard Agent can build one where a moment has no usable audio.

What if the room is noisy throughout? Then keeping a moment keeps the noise; Vizard Agent will say so.

Can it match the levels across several clips? Yes. Ask Vizard Agent to level them as a set.

Will it remove my speech in the footage? Only if you ask. Tell Vizard Agent whether the on-camera talking stays.

Can I hear the kept moments alone? Yes. Ask Vizard Agent to export them as their own track.

Does this work for a workshop or a kitchen? Yes. Vizard Agent handles any footage where specific actions make the sounds.

Will the video change? No. Vizard Agent works on the audio only.

Can it duck rather than mute? Yes, and telling Vizard Agent which you want saves a round.

Why name the action rather than the sound? Because an action happens at a time you can find, and "the good sounds" is a description of how you felt about them.