Vizard Agent

How to turn your own artwork clips into a commentary short

Last updated 2026-08-16 · 6 min read

Upload your artwork clips to Vizard Agent and say what the commentary should argue. Vizard Agent inspects every clip as a frame grid, transcribes the speech in all of them at once, checks whether anything was missed, and builds bold animated captions timed to the words actually spoken.

What is the short version?

If you make artwork and record yourself talking over it, the material for a strong short already exists. What turns it into a post is captions that match the speech exactly, because these videos are watched without sound and read rather than heard.

  1. Go to Vizard Agent and upload the clips.
  2. Say what the commentary is about and what it argues.
  3. Say the caption style and where it is going.

What do you need before you start?

Your clips and your angle. Vizard Agent transcribes and reviews everything you send, so what it needs from you is the argument — the point the short is making — since a commentary video is a position rather than a compilation.

What do you type into Vizard Agent?

Say what the clips are and what the short should argue. Vizard Agent watches and transcribes them itself, so the sentence that changes the result is the one describing the point — everything about the content it can find out by looking.

Prompt

These are clips of my artwork with my commentary. Turn them into a short [commentary] video with captions. [The argument.] Vertical, under a minute.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real commentary short. The end of the run is the interesting part: it went back to check whether anything had gone unsubtitled, found twelve quiet seconds, and fixed it by hand.

  1. Probes every clip and extracts sample frames, tiling all eight into one grid.
  2. Looks at that grid to understand what each clip shows.
  3. Checks which clips have audio streams and what codecs they use.
  4. Transcribes the speech in all of them in parallel.
  5. Compares frames from the long clips at different times to see whether they actually differ or are static.
  6. Analyses one clip's audio track for music, voice quality and levels.
  7. Trims the final twelve seconds, uploads it and transcribes it separately, because the main pass had missed quiet speech there.
  8. Extends the word-level timing file with those manually timed events.
  9. Regenerates the animated captions from the corrected file, renders, and checks the end frames to confirm the new lines appear.

Steps 7 and 8 are the ones worth reading twice. Vizard Agent noticed that a quiet passage at the end had produced no captions, isolated that audio, transcribed it on its own, and edited the timing file so the finished video subtitles every word that was spoken.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 56.94 seconds, AAC audio. Vertical, just under a minute, your own clips in sequence with bold animated captions carrying an active-word highlight, and the final quiet passage subtitled along with the rest.

Vizard Agent extracted frames from the end of the render specifically to confirm the appended lines were there, rather than assuming the rebuild had worked.

A minute is roughly eight clips, which means each one holds for about seven seconds. That is long enough for a point to land and short enough that a static shot of artwork does not start to feel like a still image.

It is not instant, and this category was not separately measured. Comparable work runs a median of 28 to 38 minutes end to end. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.

When does this not work well?

Commentary is an argument, and Vizard Agent presents whatever argument you give it with the same confidence. On political or contested subjects that is worth thinking about before the video exists rather than after it has been posted.

How do you fix a result that came back wrong?

Say what is wrong. Vizard Agent keeps the frame grids, every transcript, the corrected timing file and the caption styling separately, so a caption fix or a reorder is a re-render rather than a repeat of the whole transcription pass.

How does Vizard Agent compare to doing it yourself?

By hand this is transcribing eight clips, typing the captions out, timing them word by word, and then never noticing that the last twelve seconds carry no subtitles at all — because you already know what you said there, and your own memory fills the gap your viewers will actually see.

By hand Vizard Agent
Knowing what is in each clip Watch them all Frame grid plus parallel transcripts
Captions Type and time by hand Generated from word timings
Missed quiet speech Never noticed Isolated, re-transcribed, appended
Restyling Redo the caption pass Say the style

Common questions

Do the clips need to have my voice on them? No. Vizard Agent checks which have audio and can add a narrated line where the clips are silent.

Will every word be captioned? That is the intent. Vizard Agent checks for passages the transcription missed and fills them rather than shipping a gap.

Can it reorder my clips? Yes. Say the argument and Vizard Agent sequences them to make it, or name the order yourself.

Can I choose the caption font? Yes. Name the look — heavy, condensed, highlighted — and Vizard Agent styles them from the installed fonts.

How long should a commentary short be? Under a minute. If the argument needs longer, it needs a different format.

Can it match my previous posts? Yes. Keep them in one project and Vizard Agent reuses the caption styling across the series.

Can it add music underneath? Yes. Vizard Agent searches a licensable library and mixes the track under the commentary rather than over it.

Does it need my clips to have the same format? No. Vizard Agent probes each one and normalises them during the assembly.