Vizard Agent

How to make a premium vertical music clip from a song excerpt

Last updated 2026-08-24 · 7 min read

Upload a short excerpt of the song with reference photos and tell Vizard Agent it is for vertical feeds. Vizard Agent maps the excerpt's energy, generates cinematic shots in the referenced look, checks the hands and faces frame by frame, and places the titles where they read against the sky.

What is the short version?

Fifteen seconds of a song is enough to sell it if the pictures are good enough to stop a scroll. That means three or four generated shots that hold up to being looked at closely, and titles that survive being placed over a bright sky.

  1. Go to Vizard Agent and upload the song excerpt.
  2. Attach reference photos for the look you want.
  3. Say it is vertical, for Reels and Shorts, and how long.

What do you need before you start?

The excerpt and some references. Vizard Agent generates the shots itself, so you need no footage at all, and the reference photos are doing the heavy lifting because they define a look far more precisely than any description would.

What do you type into Vizard Agent?

Attach the references and name the format. Vizard Agent reads the photos for the look and analyses the audio for the energy, then builds shots that fit both, so your instruction mostly needs to say where the clip is going and how long it should run.

Prompt

Create a premium cinematic vertical music clip using this song excerpt and these visual references. [9:16] for [Reels and Shorts], around [15] seconds, with an opening and closing title.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real music clip. The part that consumed the most attention was the closing title, which was rebuilt four times because it kept losing legibility against a gold sunset.

  1. Reviews the reference photos and the previous release for the established look.
  2. Probes the audio and views its waveform as an energy map.
  3. Analyses the song's mood, beats and loudness to plan where the shots change.
  4. Generates three cinematic shots in parallel and retries the one that failed.
  5. Reviews the motion across frames in each shot rather than judging a single still.
  6. Checks the faces, the hands and the anatomy mid-clip in every generated shot.
  7. Renders the opening and closing titles and inspects them on dark and on gold backgrounds separately.
  8. Assembles the cut, then re-cuts with delayed titles and a cleaner fade-through-black transition.
  9. Rebuilds the closing title three more times — repositioned into the sky, recoloured to ivory, then tightened above the sun — previewing each on the real footage before rendering the final Reels cut.

Step 6 is the check that keeps a generated clip publishable. Hands and faces are where generation fails most visibly, and looking at them mid-shot rather than at the first frame is what catches a problem before it reaches an audience.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 14.37 seconds, AAC audio. Vertical, just over fourteen seconds, three generated cinematic shots cut to the song's energy, a fade through black, and opening and closing titles placed to stay legible.

Fourteen seconds carrying three shots means each one holds for around four seconds, which is long for a feed and right for this. A premium music clip asks to be looked at rather than scrolled past, and short cuts would work against that.

When does this not work well?

Generated people are the risk in this format, because a music clip invites the viewer to look closely at exactly the things generation gets wrong. Vizard Agent checks them deliberately and some shots still have to be discarded.

How do you fix a result that came back wrong?

Name the shot or the title. Vizard Agent keeps the reference analysis, every generated clip, the frame checks and all the title versions, so a replacement shot or a repositioned card is a re-render rather than a fresh build.

How does Vizard Agent compare to doing it yourself?

By hand this is generating shots and picking between them from thumbnails, missing a six-fingered hand until a commenter points it out, and dropping a white title onto a sunset where it simply disappears into the sky. Vizard Agent previews every version on the real frame instead.

By hand Vizard Agent
Judging a generated shot From one still Motion and anatomy checked across frames
Where the cuts land By feel Against the song's measured energy map
Title legibility Judge on a dark editor Previewed on dark and gold separately
A failed shot Live with it Retried, then re-checked

Common questions

Why only fifteen seconds? Because this format is a teaser for a feed, not a music video. Fifteen seconds is long enough for three shots that each hold, and short enough that Vizard Agent can spend the effort on making every one of them stand up to a close look rather than spreading it thin.

How long an excerpt should I upload? Around fifteen seconds. That is the length this format delivers.

Do I need any footage? No. Vizard Agent generates every shot from your reference photos and the song's own energy.

Will the people look real? Close, and hands and faces are the weak point. Vizard Agent checks them mid-shot before using a clip.

Can it match my previous release? Yes. Give Vizard Agent the earlier video and it studies it for continuity.

Why did the title take four attempts? Because it sits over a gold sunset, and white text on bright sky is the hardest legibility case there is. Vizard Agent previewed each version on the real frame until one worked.

Can the artist appear? Not as a generated person. If the artist must be in it, film them.

Does it cut to the beat? It cuts to the song's energy map, which for a slow cinematic clip matters more than the beat grid.

Does Vizard Agent measure the audio? Yes. Vizard Agent checks the final waveform and the balance before the clip is delivered.

Does Vizard Agent watch the final cut? Yes. It checks the motion, the audio and every text frame before delivery.