How to cut one generated clip into a cinematic reel with a voiceover
Upload your generated clip to Vizard Agent and paste the voiceover text. Vizard Agent studies the footage frame by frame to find its transitions, records two narration takes, compares their timings and tone, then cuts the picture to the meaning and rhythm of the line that is being spoken over it.
What is the short version?
A generated clip is usually one continuous piece of footage with a couple of natural turns in it. Turning it into a reel means finding those turns and making the narration land on them, rather than laying a voice over the top and hoping.
- Go to Vizard Agent and upload the clip.
- Paste the voiceover text exactly as it should be read.
- Say the length, the shape and the register of the read.
What do you need before you start?
The footage and the words. Vizard Agent works out where the clip changes and where the narration should sit against it, so the input that matters is the text — this format is a piece of writing with a picture under it.
- The clip. Generated or filmed; one piece is the normal case here.
- The voiceover text. Word for word. Vizard Agent reads it as written.
- The register. Cinematic and documentary is the convention.
- The length. Sixteen seconds is a working size when the footage is short.
- Whether the footage can be re-timed. Say if slowing or holding a section is acceptable.
What do you type into Vizard Agent?
Paste the text and name the length. Vizard Agent studies the footage itself to find its transitions, so the brief is mostly the words plus one sentence about how they should be delivered and how long the finished piece must run.
Prompt
Variants worth knowing:
- Two voices to compare. Ask for both; hearing them against your words is the fastest way to choose.
- Music-led. A version with no voice at all, cut to the track, is worth seeing side by side.
- Longer. Say so and Vizard Agent holds shots rather than speeding the read.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real cinematic reel. The middle of the run is spent on the voice rather than the picture — two takes, both transcribed, both compared, before a single cut is made.
- Inspects the uploaded video's format and duration.
- Extracts sample frames and tiles them into a montage to understand the footage.
- Checks the frame counts and rate, then samples frames around specific timestamps.
- Inspects the transition frames closely to find exactly where the footage turns.
- Extracts frames from the second half and checks its visual progression.
- Searches for fitting voices, then for cinematic documentary voices specifically.
- Generates two voiceover options and atmospheric music in parallel, and checks their durations.
- Transcribes both takes for word timings and prints them out to compare.
- Analyses the tone and quality of both options and the music, then tests the timing against the sixteen-second target.
Step 8 is the discipline. Two takes of the same script have different word timings, and printing both is how you choose the one whose rhythm actually fits the footage rather than the one that sounds better in isolation.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 16.0 seconds, AAC audio. Vertical and exactly sixteen seconds, your own footage re-cut to the meaning of the narration, with the chosen take over it and atmospheric music mixed underneath the whole thing.
Vizard Agent tested the narration's timing against the target length before assembling anything, rather than cutting the picture first and discovering afterwards that the voice did not fit inside it.
Vizard Agent delivers the chosen take's word timings alongside the cut, so a later change to the captions or to one shot's length works from numbers that already exist rather than from a fresh transcription of the same narration.
How long it takes was not measured separately here. Comparable work runs a median of 28 to 38 minutes end to end, and the upper quartile is several times the median. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.
When does this not work well?
One clip is one clip, and no amount of editing makes it two. Vizard Agent finds the real transitions inside your footage and cuts the narration to them, and a sixteen-second read laid over five seconds of usable picture means the same shot has to appear more than once.
- Short footage limits the cut. Vizard Agent can hold and re-time shots, and repetition eventually shows.
- Generated footage has few real turns. The transitions it finds are the ones that exist.
- The text carries it. A cinematic read over a thin line is an expensive way to say very little.
- Exact lengths constrain the read. Sixteen seconds of words is fewer than most people expect.
- Music can swamp a quiet voice. Ask Vizard Agent to sit the bed lower if the read is intimate rather than declamatory.
- A cinematic register sets expectations. It promises a certain kind of film, and the footage has to be good enough to keep that promise.
How do you fix a result that came back wrong?
Say which part is wrong. Vizard Agent keeps both narration takes, the word timings it printed for each, the transition analysis of your footage and the music as separate pieces, so switching takes or re-cutting the picture works from measurements it already made rather than from a fresh pass over everything.
- "Use the other take." Already transcribed; the picture re-cuts to its timings.
- "Hold the opening longer." A timing change against the same read.
- "The music is too present." A mix change only.
How does Vizard Agent compare to doing it yourself?
By hand this is generating a voice, laying it over the clip, finding that it runs three seconds long, re-recording it slightly faster, and then re-cutting everything underneath it because every timing you had just moved again — which is the loop this format traps people in.
| By hand | Vizard Agent | |
|---|---|---|
| Finding the footage's turns | Scrub for them | Transition frames inspected |
| Choosing a take | Listen once | Both transcribed and compared |
| Fitting the length | Trim and hope | Timing tested against the target |
| Re-cutting | Redo the timeline | Re-cut from stored timings |
Common questions
Can it use footage I generated elsewhere? Yes. Vizard Agent treats it as the main visual source and cuts to the narration.
Will it read my text exactly? Yes. Vizard Agent narrates the words as written unless you ask for an edit.
Can I hear two voice options? Yes, and it is worth it. Vizard Agent generates and compares takes before committing.
What if my clip is shorter than the script? Vizard Agent holds and re-times shots, and a much shorter clip will repeat. More footage is the real fix.
Can I set an exact length? Yes. Vizard Agent tests the narration against your target before assembling anything.
Does it add music? Yes. Vizard Agent generates or sources an atmospheric bed and mixes it under the voice rather than over it.
Can it add subtitles? Yes, and cheaply, since the word timings already exist from the take it chose.
Can I use several clips? Yes. Vizard Agent finds the transitions in each and cuts between them as well as within them.