How to put a new voice track on a video and keep it in step
Upload the video and the new audio. Vizard Agent transcribes both tracks to compare what they say and when, maps every text card and shot change in the picture, then re-times the video so each card still arrives on the words that belong with it.
What is the short version?
Swapping one audio file for another is a one-line operation that takes seconds. Keeping the picture in step with it is the actual job, because every text card and every cut in that video was timed against a narration which no longer exists.
- Go to Vizard Agent and upload the video.
- Upload the new audio track.
- Say what has to line up — the text cards, the shot changes, or both.
What do you need before you start?
Both files, and an idea of how closely the new script follows the old one. Vizard Agent compares the two transcripts, so a new recording that says roughly the same things in the same order needs far less re-timing than one that has been rewritten.
- The video. With its original audio still on it.
- The new audio. As recorded, uncut.
- The relationship. Same script, or a rewrite.
- What must land. Cards, cuts, or specific moments.
- The ending. Whether the video trims to the audio.
What do you type into Vizard Agent?
Say explicitly what has to stay aligned once the audio changes. Vizard Agent will otherwise simply attach the new track and hand it back, which is technically exactly what you asked for and leaves every text card arriving at the wrong moment.
Prompt
Variants worth knowing:
- "Keep the cards on the right words." The line that matters.
- Say whether it is a rewrite. It changes how much re-timing is needed.
- Say what to do with the ending. Trim, hold, or extend.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real audio swap for a slide-based video. Most of the session happens before the swap itself: it is spent working out where the picture's existing timing came from, because that is the thing the new track has to fit into.
- Checks the duration and format of both the video and the new audio.
- Extracts sample frames from the video and reviews them.
- Analyses the new audio and the video's original audio.
- Transcribes both tracks to compare their content and their timing.
- Extracts a frame per second to analyse the video's slide structure.
- Analyses the shot and slide changes across the video.
- Builds thumbnails of every shot and reviews them.
- Checks the word-by-word transcript in detail.
- Analyses the text and the duration of every text card in the video.
- Checks the new audio's exact duration and its ending.
- Measures the new audio's loudness and reads the full values.
- Combines the video with the new audio and trims to its length.
- Extracts the closing frames and builds a montage of them to review.
- Uploads the result and analyses the finished video end to end.
Step nine is the step that separates this from an attach-and-hope. Knowing what each card says and how long it holds is what lets the picture be re-timed around a narration that arrives at different moments.
What does the result look like?
The video this page is written from came back with the new voice track in place, the picture trimmed to match its length, the text cards still arriving with the words they belong to, and the audio level measured rather than left wherever the recording happened to sit.
The closing frames were checked separately. A swapped audio track almost always ends at a different moment from the original, and the end of a video is where that shows most.
When does this not work well?
Re-timing a picture around a new voice track has real limits, and almost all of them come from how far the new script has drifted away from the old one. These are the situations where something has to give.
- A complete rewrite. There may be nothing to align to.
- Much longer audio. The picture runs out before the voice does.
- Much shorter audio. Cards get cut before they can be read.
- Burnt-in captions. They show the old words regardless.
- Music tied to the old cut. It has to be rebuilt too.
How do you fix a result that came back wrong?
Name the card that is landing wrong, or the moment in the video. Vizard Agent keeps both transcripts, the full shot map and every measured card timing, so a re-aligned section comes back as a straight re-render rather than needing another round of analysis from the top.
- "That card comes up too early." Re-timed against the new transcript.
- "The ending is abrupt." Held longer, or the last shot extended.
- "The voice is quiet." Re-levelled to a measured target.
How does Vizard Agent compare to doing it yourself?
By hand the swap itself takes seconds and fixing everything it broke takes the rest of the afternoon: every card nudged into place, every cut re-checked against the new voice, and the ending rebuilt from scratch. Vizard Agent compares the two transcripts first and re-times the picture from that.
| By hand | Vizard Agent | |
|---|---|---|
| The swap | Instant | Instant |
| The cards | Nudged one at a time | Re-timed from the new transcript |
| Knowing what changed | Watched to find out | Both tracks transcribed and compared |
| The ending | Discovered on export | Checked as its own step |
Common questions
Can it just attach the audio? Yes, if that is all you want. Vizard Agent will say when the cards will end up misaligned.
What if my new script is different? It compares both transcripts. Vizard Agent tells you how far apart they are.
Will it trim the video? Yes, to the audio's length, if you ask. Vizard Agent checks the closing frames afterwards.
What if the new audio is longer? Something has to give. Vizard Agent can hold shots longer or extend the last one.
Does it fix the loudness? Yes. Vizard Agent measures the new track and levels it rather than leaving it as recorded.
What about burnt-in captions? They will show the old wording. Vizard Agent can remove and rebuild them as a separate pass.
Can it keep the original sound effects? Yes, if they are separable. Vizard Agent can split them from the old narration.
Does it work for a different language? Yes. Vizard Agent compares timing rather than meaning when the languages differ.
Can I use my own voice? That is the usual case. Vizard Agent takes whatever recording you upload.
What about the music? It may need rebuilding. Vizard Agent flags when the music was cut to the old timing.
Why transcribe the old audio at all? Because it tells you why the picture is timed the way it is. Without it, every card is a guess about which words it was waiting for.
Can it re-time the shot changes too? Yes. Vizard Agent maps the cuts as well as the cards and can move both.
Does it check the result? Yes. Vizard Agent analyses the finished video end to end after rendering.