Vizard Agent

How to fix uneven levels without the voice sounding processed

Last updated 2026-09-22 · 8 min read

Give Vizard Agent the levels you need and say the voice must still sound like the person. Levelling works by pushing quiet parts up and loud parts down, and pushed far enough that makes any voice sound gritty — so the check that matters is a comparison against the original, not a meter reading.

What is the short version?

You asked for even levels and Vizard Agent delivered them, and now the voice sounds harder and grainier than the person actually does. Both things are true at once: the numbers are right and the recording is worse than it was.

  1. Tell Vizard Agent the floor and ceiling you need.
  2. Say the voice must stay natural and organic, like the source.
  3. Ask it to compare the processed voice against the original.

What do you need before you start?

The recording and the numbers. Vizard Agent works to stated targets — a peak ceiling, a floor for the quiet passages — and having those written down is what lets it tell you when they are in conflict with keeping the voice natural.

What do you type into Vizard Agent?

Say "cosmetic" and mean it literally. Asking Vizard Agent for professional-sounding audio invites the entire processing chain — compression, de-essing, an exciter — when what you actually want is the very same voice with only its level problems taken out of it.

Prompt

Level out the audio on this source video, maximum -6dB, and fix any mic pops. Do not modify or change the voice itself — only cosmetic audio fixes so it sounds professional. The voice should still sound natural and organic, exactly like the source, just balanced.

Variants worth knowing:

What does Vizard Agent actually do?

Here is a stretch of the Vizard Agent sequence on a talking-head clip levelled to stated targets. The interesting part is where the levelling actually gets applied: to the quiet passages specifically rather than across the whole file at one setting.

  1. Renders a revision with broader cuts and more room at the speech boundaries.
  2. Checks every speech join and the last seconds for clipped words.
  3. Finds the ending leaves the story hanging and re-cuts to a natural close.
  4. Checks the final waveform and the quiet space after the last word.
  5. Selects complete spoken phrases for the hook, the explanation and the resolution.
  6. Confirms the audio stays below the peak limit and the last word has room to finish.
  7. Measures the quiet speech passages and checks the exact word boundary before a join.
  8. Checks a phrase boundary so removing one word does not clip the start of the next.
  9. Applies speech-aware levelling, with extra attention to the quiet passages.
  10. Checks the spoken peaks after the quiet speech has been lifted.
  11. Reviews the export for natural voice tone, even loudness and a smooth handoff.
  12. Compares the before-and-after waveform, including the lifted passages and the controlled peaks.
  13. Compares the original voice with the processed audio and judges how much peak control is actually needed.

Step thirteen is the one that answers the complaint. The test is not whether the file meets the target; it is whether the voice still sounds like the person when you play the two back to back — and if it does not, the processing comes down rather than the target changing.

Step nine is why it can come down. Levelling applied to the whole file has to be aggressive enough for the quietest moment; levelling applied where the speech is quiet lets everything else be left alone.

Step eight is a detail that matters more than it sounds. Tightening a join by removing a word can clip the start of the next one, and a clipped consonant is heard as bad audio even when the levels are perfect.

What does the result look like?

Audio from Vizard Agent that sits inside your stated targets and still sounds like the person: the quiet passages audible, the peaks controlled, no grit or hardness on the voice, and nothing clipped at the joins between phrases.

When does this not work well?

Some recordings cannot be levelled gently. If the original is clipped, recorded far from the microphone, or has a noise floor close to the speech, the processing needed to make it even is the processing that makes it sound processed.

How do you fix a result that came back wrong?

Describe the sound rather than the setting you think caused it. "Gritty", "hard", "boxy" and "thin" each point at a different stage of processing, and Vizard Agent can ease that one stage rather than reducing everything at once.

How does Vizard Agent compare to doing it yourself?

By hand you run a loudness normaliser and a limiter, which is one click each and makes all of the numbers correct. Whether the result still sounds like the person is a separate question entirely, and the tools themselves never ask it.

By hand Vizard Agent
The target A number A number plus the source as reference
Where applied The whole file The passages that need it
Verification The meter The original voice, compared
Joins Trimmed tight Checked for clipped words
Result Even and processed Even and recognisably the person

Common questions

What targets should I ask for? If you have platform requirements, use those; otherwise Vizard Agent will suggest a range.

Can it remove mic pops? Yes, and that is a cosmetic fix rather than a change to the voice.

Will it use compression? Vizard Agent uses as little as it can when you ask for a natural voice.

Can it match another video's loudness? Yes. Give Vizard Agent the reference and it matches it.

What about music underneath? Levelled separately, so the voice does not have to fight it.

Does it change the picture? No. Vizard Agent treats this as an audio-only pass.

Can it fix one section only? Yes, and Vizard Agent finds that gentler than treating the whole file.

Will the file be quieter overall? Not necessarily — controlled peaks often allow a higher average.

What if two speakers sound different? Vizard Agent can level them separately and match them to each other.

Can I hear a before and after? Yes. Ask Vizard Agent for both and listen on headphones.

Is a re-record better? For clipped audio, yes. Vizard Agent will tell you when it is.

Can it hit a platform's loudness standard? Yes, without flattening the voice to get there.

What about breath and mouth noise? Reducible, though removing all of it is what makes a voice sound artificial.

How do I check it? Play the original and the processed version back to back on the same speakers.