Vizard Agent

How to make dark cinematic shorts with big text on screen

Last updated 2026-08-16 · 6 min read

Give Vizard Agent your script and your spec — the voice, the music level, how many words appear at once. Vizard Agent auditions a narrator against that description first, records every segment, searches stock footage for the specific shot each line calls for, and builds the text to your rules.

What is the short version?

This format is a spec more than a creative brief: a low calm voice, ambient music kept well under it, one to three words on screen at a time, and a shot that literally illustrates each line. Get the spec right and the video makes itself.

  1. Go to Vizard Agent and give it the script or the theme.
  2. Describe the narrator: age, register, pace.
  3. Give the text rules and the music level you want.

What do you need before you start?

The script and the spec. This is one of the few formats where naming numbers helps — words per screen, music volume, segment lengths — because the whole look comes from restraint applied consistently rather than from any individual choice.

What do you type into Vizard Agent?

Write the spec as a spec rather than as a description. Vizard Agent reads a list of constraints as constraints and holds to them, so numbers, colours and pacing rules do far more work here than any amount of writing about the mood or the feeling you are hoping to land.

Prompt

Voice: [low, calm, male, measured]. Music: [dark ambient], kept at [10-15%]. Text: large, [1-3] words on screen, [white with a yellow accent]. Length: [40] seconds. Script: [yours, or the theme].

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real set of cinematic shorts. The voice comes first and everything else is built to it — including the footage search, which is done line by line rather than as one mood board.

  1. Checks the voice, music and video tools and their costs.
  2. Searches for a narrator matching your description in the script's language.
  3. Generates a test segment to hear the voice before committing.
  4. Tries an alternative voice when the first is not right.
  5. Checks the first segment's duration, then generates the remaining four in parallel.
  6. Checks the duration of every segment so the visuals can be built to fit.
  7. Searches the music library for the exact genre named in the spec and downloads it.
  8. Searches stock footage segment by segment — a man on the edge of a bed, a sleeping face, a clock.
  9. Keeps searching per line until every segment has the specific shot its words describe.

Step 8 is the difference between this format working and not. Vizard Agent searched for the literal image each line described rather than assembling a general mood, which is why the text and the picture agree with each other.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 41.0 seconds, AAC audio. Vertical, forty-one seconds, five narrated segments over sourced footage with large text on screen and the music sitting well under the voice.

If the voice is hard to hear, say so — Vizard Agent lowers the music against the stems rather than re-rendering the video, which is a fast change.

Timing was not measured separately for this kind of job. The closest measured work runs a median of 28 to 38 minutes end to end, with the middle half spread considerably wider. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.

When does this not work well?

The format is saturated and its production is easy, which puts all the weight on the script. Vizard Agent will make any script look like this, and a video that looks the part while saying nothing is the most common outcome in this genre.

How do you fix a result that came back wrong?

Name the segment or the setting that is wrong. Vizard Agent keeps every narrated segment, its measured duration, the chosen music and each downloaded clip as separate pieces, so lowering the music level or swapping a single shot leaves the read, the text and the timing of everything else exactly untouched.

How does Vizard Agent compare to doing it yourself?

By hand this is a text-to-speech tool, a stock site, a music track and an editor — five segments, five searches, five text animations, and a mix. It is a couple of hours per short, on a format that only works if you post many.

By hand Vizard Agent
Casting the voice Scroll a voice list Auditioned against your spec
Footage per line Search one at a time Searched per segment, in order
Timing the text Nudge each caption Built to the measured segments
The mix Re-render to adjust Stems kept, adjust and re-mix

Common questions

Can it write the script? Yes, from a theme. Your own writing is what separates one of these from the thousands that look identical.

Can it use my voice? Yes. Provide a clean sample and Vizard Agent clones it for the narration.

Can it work in my language? Yes. Vizard Agent searches for a narrator in that language and generates the read there.

How loud should the music be? Well under the voice. Name a level and Vizard Agent mixes to it rather than guessing.

How long should one be? Around forty seconds, which is about five segments. Longer loses the tension the format depends on.

Can I make a series with the same voice? Yes. Keep them in one project and Vizard Agent reuses the same narrator and rules.

Can it use my own footage? Yes, and it helps. Vizard Agent prefers your material and searches stock only for the lines you have not covered.

How many segments fit in forty seconds? Around five. Vizard Agent measures each narrated segment and builds the visuals to those durations rather than to even splits.

Can the text be a different colour scheme? Yes. Name the colours and what each one is for, and Vizard Agent applies that rule across every segment.