How to change the voice and music in a reel you already made
Give Vizard Agent the reel and say what the voice and music should become. Vizard Agent generates the new narration in the register you name, sources matching music, rebuilds the scenes with the graphics measured into place, and adds scrims behind the text where it was fighting the picture.
What is the short version?
Changing the voice on a finished reel is not a swap, it is a rebuild. The new narration has a different length and different pauses, so every scene has to be re-timed to it, which is the moment to fix the things that were not quite right the first time.
- Go to Vizard Agent and give it the reel and its source photos or clips.
- Say what the voice should become and in which language.
- Say the music style and anything else that should change.
What do you need before you start?
The reel and its raw material. Vizard Agent can work from the finished file alone and does much better with the original photos and clips, because rebuilding the scenes to a new narration length means re-cutting rather than stretching what is already rendered.
- The existing reel. So the structure is clear.
- The source photos and clips. Ideally the originals.
- The new voice. Deeper, warmer, older, a different language.
- The music style. Worship, ambient, upbeat, cinematic.
- The script. If the words are changing too.
What do you type into Vizard Agent?
Describe the voice as a person would. Vizard Agent searches by that description rather than by settings, so "a deep Hindi voice" gets you a much better shortlist than any parameter, and naming the music style in the same sentence keeps the two matched to each other.
Prompt
Variants worth knowing:
- A second language version. Vizard Agent can output the same reel with captions in another language.
- Smoother transitions. Say so; it is a common complaint about a first version.
- Scrims behind the text. They fix captions that are fighting a busy photo.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real rebuild. The step people skip is second: it checked the photos' orientation metadata, because a photo that displays upright in a phone gallery can be stored rotated and will come out sideways in a render.
- Downloads the reel and the source photos and builds a contact sheet of them.
- Checks the photos' orientation metadata and re-reviews them the right way up.
- Searches for a deep voice in the target language and generates worship-style music.
- Generates the new narration and the title graphics, then assembles the voice track and measures its level.
- Renders the graphics over a background and inspects each one frame by frame at full resolution.
- Prepares vertical versions of every photo and renders each scene with gentle motion.
- Re-renders the scenes at the right length once the new narration's timing is known.
- Generates captions, renders the full reel, measures the music slice and re-renders with the level corrected.
- Measures where each graphic sits in frame, reframes the landscape photos, builds soft scrims behind the text, then produces a second version with captions in another language.
Step 9 is the fix that makes the difference visually. Text over a photograph is legible or not depending on what is behind it, and a soft scrim costs nothing while turning an unreadable caption into a readable one.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 61.0 seconds, AAC audio. Vertical, sixty-one seconds, rebuilt around a new deep-voiced narration with worship music, gentle motion on every photo, scrims behind the text, and a second version delivered with captions in another language.
Two language versions from one rebuild is the efficiency here. The scenes, the timing and the graphics are shared; only the caption track changes, so the second version costs a render rather than a second job.
When does this not work well?
Rebuilding around a new voice means the timing changes everywhere, and anything that was locked to the old read has to move with it. Vizard Agent handles the re-timing and some source material limits what can be done.
- A new voice changes the length. Every scene re-times, so this is a rebuild rather than an audio swap.
- Without the originals, options narrow. Working from the rendered reel alone means no reframing.
- Landscape photos in a vertical reel. They need reframing or a scrim, and something gets lost.
- Religious and cultural material deserves care. Get the wording reviewed by someone who knows the context.
- Burnt-in text from the old version stays. Unless the source files come with it.
How do you fix a result that came back wrong?
Name the scene or the line. Vizard Agent keeps the new narration with word timings, the graphics measurements, the reframed photos and both caption tracks, so a change of wording or a caption restyle is a re-render rather than another rebuild.
- "The voice is not deep enough." Another candidate, generated and measured.
- "The caption sits on her face." Repositioned against the measured frame coordinates.
- "The transition is abrupt." Re-rendered with a longer move.
How does Vizard Agent compare to doing it yourself?
By hand this is regenerating a voiceover, discovering it runs eight seconds longer than the one it replaces, re-timing every single scene to match the new read, and then finding that three of your photos have come out sideways because of their rotation metadata.
| By hand | Vizard Agent | |
|---|---|---|
| A longer new narration | Stretch or trim the scenes | Every scene re-rendered to the new timing |
| Photo orientation | Find out at render | Metadata checked before anything is built |
| Text legibility | Judge on a monitor | Graphics measured, scrims added where needed |
| A second language | Rebuild the reel | Caption track swapped, everything else shared |
Common questions
Can it just replace the audio? Not cleanly. A new voice has a different length and different pauses, so the scenes have to be rebuilt around it.
Do I need the original photos? Not strictly, and Vizard Agent does much better with them. Reframing and re-timing both need the source files.
Can I change the language too? Yes. Vizard Agent searches for a voice in the new language and builds matching captions.
Will it fix the transitions? Yes if you ask. Vizard Agent treats smoother transitions as a real instruction, and it is one of the most common second-version requests.
Why do captions need a scrim? Because text over a photograph is only as legible as what is behind it. A soft darkening behind the line fixes it without hiding the picture.
Can I get two language versions? Yes, and the second is cheap since the scenes and timing are shared.
Will it check the photo orientation? Yes. Vizard Agent reads the metadata rather than trusting how the file previews.
Does Vizard Agent measure the mix? Yes. Vizard Agent measures the voice and the music separately and re-rendered this reel once after finding the music sitting too high against the narration.
Does Vizard Agent check the finished reel? Yes. It pulls key frames at full detail and checks the title, the times and the closing card for legibility.