Vizard Agent

How to cut one generated clip into a cinematic reel with a voiceover

Last updated 2026-08-17 · 6 min read

Upload your generated clip to Vizard Agent and paste the voiceover text. Vizard Agent studies the footage frame by frame to find its transitions, records two narration takes, compares their timings and tone, then cuts the picture to the meaning and rhythm of the line that is being spoken over it.

What is the short version?

A generated clip is usually one continuous piece of footage with a couple of natural turns in it. Turning it into a reel means finding those turns and making the narration land on them, rather than laying a voice over the top and hoping.

  1. Go to Vizard Agent and upload the clip.
  2. Paste the voiceover text exactly as it should be read.
  3. Say the length, the shape and the register of the read.

What do you need before you start?

The footage and the words. Vizard Agent works out where the clip changes and where the narration should sit against it, so the input that matters is the text — this format is a piece of writing with a picture under it.

What do you type into Vizard Agent?

Paste the text and name the length. Vizard Agent studies the footage itself to find its transitions, so the brief is mostly the words plus one sentence about how they should be delivered and how long the finished piece must run.

Prompt

Create a cinematic [16] second reel using my uploaded video and this voiceover text. Use the video as the main visual source and cut it to match the meaning and rhythm of the words. [paste the text]

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real cinematic reel. The middle of the run is spent on the voice rather than the picture — two takes, both transcribed, both compared, before a single cut is made.

  1. Inspects the uploaded video's format and duration.
  2. Extracts sample frames and tiles them into a montage to understand the footage.
  3. Checks the frame counts and rate, then samples frames around specific timestamps.
  4. Inspects the transition frames closely to find exactly where the footage turns.
  5. Extracts frames from the second half and checks its visual progression.
  6. Searches for fitting voices, then for cinematic documentary voices specifically.
  7. Generates two voiceover options and atmospheric music in parallel, and checks their durations.
  8. Transcribes both takes for word timings and prints them out to compare.
  9. Analyses the tone and quality of both options and the music, then tests the timing against the sixteen-second target.

Step 8 is the discipline. Two takes of the same script have different word timings, and printing both is how you choose the one whose rhythm actually fits the footage rather than the one that sounds better in isolation.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 16.0 seconds, AAC audio. Vertical and exactly sixteen seconds, your own footage re-cut to the meaning of the narration, with the chosen take over it and atmospheric music mixed underneath the whole thing.

Vizard Agent tested the narration's timing against the target length before assembling anything, rather than cutting the picture first and discovering afterwards that the voice did not fit inside it.

Vizard Agent delivers the chosen take's word timings alongside the cut, so a later change to the captions or to one shot's length works from numbers that already exist rather than from a fresh transcription of the same narration.

How long it takes was not measured separately here. Comparable work runs a median of 28 to 38 minutes end to end, and the upper quartile is several times the median. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.

When does this not work well?

One clip is one clip, and no amount of editing makes it two. Vizard Agent finds the real transitions inside your footage and cuts the narration to them, and a sixteen-second read laid over five seconds of usable picture means the same shot has to appear more than once.

How do you fix a result that came back wrong?

Say which part is wrong. Vizard Agent keeps both narration takes, the word timings it printed for each, the transition analysis of your footage and the music as separate pieces, so switching takes or re-cutting the picture works from measurements it already made rather than from a fresh pass over everything.

How does Vizard Agent compare to doing it yourself?

By hand this is generating a voice, laying it over the clip, finding that it runs three seconds long, re-recording it slightly faster, and then re-cutting everything underneath it because every timing you had just moved again — which is the loop this format traps people in.

By hand Vizard Agent
Finding the footage's turns Scrub for them Transition frames inspected
Choosing a take Listen once Both transcribed and compared
Fitting the length Trim and hope Timing tested against the target
Re-cutting Redo the timeline Re-cut from stored timings

Common questions

Can it use footage I generated elsewhere? Yes. Vizard Agent treats it as the main visual source and cuts to the narration.

Will it read my text exactly? Yes. Vizard Agent narrates the words as written unless you ask for an edit.

Can I hear two voice options? Yes, and it is worth it. Vizard Agent generates and compares takes before committing.

What if my clip is shorter than the script? Vizard Agent holds and re-times shots, and a much shorter clip will repeat. More footage is the real fix.

Can I set an exact length? Yes. Vizard Agent tests the narration against your target before assembling anything.

Does it add music? Yes. Vizard Agent generates or sources an atmospheric bed and mixes it under the voice rather than over it.

Can it add subtitles? Yes, and cheaply, since the word timings already exist from the take it chose.

Can I use several clips? Yes. Vizard Agent finds the transitions in each and cuts between them as well as within them.