How to keep the audio and replace every frame of a video
Upload the video and tell Vizard Agent to keep the audio and replace the picture. Vizard Agent transcribes the narration, measures its loudness, generates an entirely new set of shots timed to what is being said, adds captions in the right typeface, and delivers a new film on the old soundtrack.
What is the short version?
Sometimes the recording is good and the video is not. The words work, the delivery works, and what lets it down is that someone spoke into a phone in a badly lit room — which is a picture problem, and the picture is the part that can be replaced.
- Go to Vizard Agent and upload the video.
- Say to keep the original audio and build new visuals for it.
- Say the mood the new footage should have.
What do you need before you start?
The recording and a sense of the look. Vizard Agent transcribes the audio and times the new shots to it, so the audio itself is the brief, and what you add is the visual direction it cannot infer from the words alone.
- The recording. Only the audio will survive.
- The mood. Atmospheric, bright, sombre, formal.
- Any branding. A logo or a previous project to match.
- The language. It decides the captions and the typeface.
- The frame. Square, portrait, vertical or wide.
What do you type into Vizard Agent?
Say what stays and what goes, in that order. Vizard Agent needs to know the audio is fixed and everything visual is open, because that is what turns the job from an edit into a rebuild with a locked soundtrack underneath it.
Prompt
Variants worth knowing:
- Branding carried from a previous project. The logo read and reused.
- Generated shots rather than stock. When nothing filmed would fit.
- Music underneath. Generated to the full length of the narration.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real rebuild where only the audio survived. The step worth noticing is early: it measured the narration's loudness before generating anything, so the new music could be built to sit under a known level rather than fought with afterwards.
- Checks the workspace and the uploaded file, then downloads it and extracts frames.
- Transcribes the audio and looks at the original frames to understand the content.
- Checks the production tool prices and pulls the branding file from a previous project.
- Reviews the logo and checks the video generation and caption options.
- Loads the atmospheric montage playbook and checks which typefaces are available.
- Extracts the narration and measures its loudness, twice, to be certain.
- Finds and installs the typefaces the captions need, then sorts the font set out.
- Generates all the new shots and the background music, then previews every shot.
- Reviews the generated shots, regenerates the closing one until it works, and builds full-length music to fit.
Step 7 is easy to overlook and impossible to skip. Captions in a language the box has no typeface for come out as boxes or broken letterforms, and installing the right font is the difference between a finished film and an unusable one.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1350, H.264, 25fps, 84 seconds, AAC audio. The four-by-five portrait frame that Instagram gives the most space to, eighty-four seconds long, carrying the original eighty-second narration with an entirely new set of generated shots above it.
Nothing from the original picture survives. What survives is the voice, the timing and the meaning, which were the parts worth keeping in the first place.
When does this not work well?
Replacing every frame is a bigger change than it sounds, and it only works when the audio can genuinely stand on its own. Vizard Agent will build the film either way, so these are worth checking honestly before you commit a recording to it.
- The audio has to be clean. New pictures do not fix a bad recording.
- A speaker on camera disappears. If the person is the point, do not replace them.
- Generated shots approximate. They illustrate a subject rather than document it.
- Long narration needs many shots. Eighty seconds is a lot of new picture.
- Sensitive subjects need care. Religious, medical and historical material carries obligations.
How do you fix a result that came back wrong?
Name the shot or the section that is not working. Vizard Agent keeps the transcript, the loudness measurements, every generated shot and the caption timings, so a single shot can be regenerated against the same narration without rebuilding the film around it.
- "The closing shot is wrong." Regenerated on its own, as it was here.
- "The music covers the voice." Re-mixed against the measured narration level.
- "The captions break." The typeface is replaced and the captions re-rendered.
How does Vizard Agent compare to doing it yourself?
By hand this means stripping the audio, finding or shooting eighty seconds of new picture, cutting it to a voice you cannot alter, and then discovering your caption font does not support the language. Each of those is small; together they are an afternoon.
| By hand | Vizard Agent | |
|---|---|---|
| New footage | Shoot or licence it | Generated to the mood and the words |
| Timing | Cut picture to a fixed track | Shots built to the transcript |
| Mixing | Balance by ear | Music built to a measured level |
| Captions | Whatever font is installed | The right typeface installed first |
Common questions
Does the original audio change at all? No. Vizard Agent keeps the narration exactly as recorded and builds everything else around it.
Can I keep some of the original footage? Yes, if you say so. Vizard Agent replaced all of it here because that was the instruction.
How does Vizard Agent time the new shots? From the transcript, so each shot changes near a natural break in what is being said.
Will the captions work in my language? Vizard Agent checks the installed typefaces and installs what the language needs before rendering.
Can it match my existing branding? Yes. It pulled the logo from a previous project on this run rather than asking again.
Why measure the narration's loudness up front? Because the music has to be built to sit under a known level. Generating a track first and then fighting it down in the mix is how narration ends up buried under its own soundtrack.
What frame shapes can it deliver? Any shape. Vizard Agent delivered four-by-five portrait here, which suits Instagram's feed.
Are the new shots stock or generated? Either one. Vizard Agent generated them here because no filmed footage would have matched the subject.
How long can the audio be? Eighty seconds here. Longer narration simply means Vizard Agent builds more shots and more music.
Does this work for a sermon or a lecture? Yes, and it is one of the better uses for it. Vizard Agent keeps the delivery intact and replaces a static camera pointed at a lectern with pictures that follow what is being said.
Does Vizard Agent check the finished video? Yes. It reviews every generated shot and regenerates the ones that do not hold.