How to replace one line of narration in a mix you cannot unmix
Send Vizard Agent the finished mix and a clean recording of the corrected line, and say which phrase is wrong. It locates the phrase by word timing, takes out only that stretch of speech while the music and ambience continue underneath, and drops the new line into the same window.
What is the short version?
The narrator says "five months later" and it should be four. The problem is that the voice, the music, the effects and the room tone are all one flattened file, so there is no voice track to edit. Vizard Agent works on the phrase, not the track.
- Send Vizard Agent the full mix and the replacement line.
- Name the wrong phrase in words.
- Ask for a transcript of the result as proof.
What do you need before you start?
Two files. The finished mix as it stands, and a recording of the corrected line on its own — ideally the same voice, the same microphone and the same room, because the new line has to sit inside a bed that never stops. A clean take with a little silence either side gives Vizard Agent room to set the edge ramps.
You also need to know what the wrong phrase actually says, word for word. "The bit about the months" is enough for a human listening through; "five months later" is what lets Vizard Agent find the exact word boundaries in a transcript rather than hunting by ear through a long track.
What do you type into Vizard Agent?
Be explicit that the rest of the mix is untouchable. This is one of the few jobs where the negative instructions matter more than the positive one, because the obvious shortcut — regenerate the whole audio — destroys exactly what you are trying to keep.
Use the mp3 as the source. Find "five months later" and replace only that phrase with the wav. Keep the original music, effects, ambience and timing exactly as they are.
That is close to the real brief behind this article, and the emphasis was well placed. Vizard Agent treats "do not recreate the music" as a hard constraint rather than a preference, which rules out the easy route before it starts.
What does Vizard Agent actually do?
It works out where the phrase is before it touches anything. Both files get transcribed, the word-level timings give the exact start and end of the spoken phrase, and loudness is measured on both so the new line can be set to match rather than guessed at. Only then does any audio get written.
In that session the working sequence was:
- Transcribed the full mix and the replacement clip, and read the word-level timing files for the phrase boundaries.
- Measured integrated loudness on both, and generated spectrograms to see where the narration band sits against the music bed.
- Removed the original speech only inside its word window, leaving the bed running.
- Dropped the replacement line in with short edge ramps so the join does not click.
- Rendered at the exact original duration, with no codec padding.
The spectrogram step is what makes the rest work. Seeing where the voice sits in the frequency range is how Vizard Agent removes the old speech without taking a bite out of the music underneath it.
What does the result look like?
One audio file, the same length as the original, that sounds like the original except for the corrected words. The music does not dip, the ambience does not break, and the edit points are not audible. Vizard Agent delivers it at the exact source duration, because a file that is forty milliseconds longer will drift against picture.
You also get the proof. Vizard Agent transcribes the delivered file and reads back what it says, so the confirmation that the narrator now says "four" is a transcript of the actual output rather than an assurance that the right command ran.
When does this not work well?
When the replacement line is a different length than the original and the picture is locked. A shorter line leaves a hole in the narration; a longer one either overruns the next sentence or has to be sped up. Vizard Agent will tell you which way the difference falls, but a big mismatch is a real edit, not a swap.
It also struggles when the old speech and the music occupy the same frequencies at similar levels. Removing the voice then takes audible bites out of the bed. In that session the first attempts at simply suppressing the old phrase were not clean enough, and the fix was to cut the segment out explicitly and filter the old speech band rather than turn it down.
And it cannot work at all if the wrong phrase never appears in the transcript. A mumbled or heavily processed line may not transcribe reliably, and then Vizard Agent is looking for something it cannot locate — give it a timestamp instead.
How do you fix a result that came back wrong?
Listen for the old words still being faintly present. That is the commonest failure, and it means the removal was a volume reduction rather than a cut. Ask Vizard Agent to mute the original entirely across the word window instead of attenuating it, which is exactly the correction that session needed.
If the join clicks, say where. Vizard Agent lengthens the edge ramps, which costs a few milliseconds of bed at the seam and is inaudible against music. A click almost always means the cut landed on a non-zero sample rather than that anything structural is wrong.
If the new line sounds louder or closer than the rest of the narration, say so. Vizard Agent matched integrated loudness, but a line recorded on a different microphone can match in level and still not match in tone, and a small amount of shaping fixes it.
How does Vizard Agent compare to doing it yourself?
In a DAW this is a familiar job and a fiddly one. You would zoom in, find the phrase, cut it, drop the new take in and spend most of your time on the crossfades and on whether the bed survived. The hard part is not the technique, it is that you are doing it by eye on a waveform where the voice and the music look the same.
The difference is the transcript, at both ends. Vizard Agent finds the phrase by what was said rather than by scrubbing, and checks the result by what is now said rather than by listening once and believing it. On a correction this small — one word, in a file that otherwise must not change — that verification is most of the value.
Common questions
Do I have to re-record the whole narration? No. Vizard Agent needs only the corrected line as a separate file.
What if I do not have a recording of the new line? Vizard Agent can generate one, though matching an existing human narrator convincingly is much harder than matching a generated one.
Will the music dip where the edit is? It should not. Vizard Agent removes only the speech, and the bed runs continuously through the window.
Can it replace more than one phrase? Yes. Each one is located by its own word timing, so Vizard Agent handles several in the same pass.
Does the file stay the same length? Yes, if the new line fits. Vizard Agent renders at the exact source duration rather than letting the encoder pad it.
How do I know the old words are gone? Ask Vizard Agent for a transcript of the delivered file. That is a reading of the output, not a report on the process.
What about video that is already cut to this audio? Unchanged, as long as the duration holds. That is why the exact-length render matters.
Can it fix a word rather than a phrase? Yes, though a single word gives Vizard Agent a very short window to work in and the join is harder to hide.
Will the ambience sound different at the edit? Not if the bed is continuous. Vizard Agent keeps the original bed running underneath rather than rebuilding it.
What if the voices do not match? Vizard Agent will say so. A mismatch in the recording is not something a level match can hide.
Does it work on a stereo mix? Yes. Vizard Agent measures the channel balance first so the replacement sits in the same place in the image.
Can I hear the before and after? Ask for both. Vizard Agent uploads the corrected file and you still have the original.
Is the original file modified? No. Vizard Agent works on copies and delivers a new file.
What if the phrase appears twice? Say which one. Vizard Agent lists both timestamps rather than guessing.