How to make dark cinematic shorts with big text on screen
Give Vizard Agent your script and your spec — the voice, the music level, how many words appear at once. Vizard Agent auditions a narrator against that description first, records every segment, searches stock footage for the specific shot each line calls for, and builds the text to your rules.
What is the short version?
This format is a spec more than a creative brief: a low calm voice, ambient music kept well under it, one to three words on screen at a time, and a shot that literally illustrates each line. Get the spec right and the video makes itself.
- Go to Vizard Agent and give it the script or the theme.
- Describe the narrator: age, register, pace.
- Give the text rules and the music level you want.
What do you need before you start?
The script and the spec. This is one of the few formats where naming numbers helps — words per screen, music volume, segment lengths — because the whole look comes from restraint applied consistently rather than from any individual choice.
- The script, or the theme. Vizard Agent will write it; your own words are stronger if you have them.
- The voice. "Low, calm, male, 45+, measured pace" is exactly the right level of detail.
- The text rules. One to three words per screen, and which colours mean what.
- The music level. Ten to fifteen per cent under the voice is typical for this format.
- The length. Forty seconds is around five segments.
What do you type into Vizard Agent?
Write the spec as a spec rather than as a description. Vizard Agent reads a list of constraints as constraints and holds to them, so numbers, colours and pacing rules do far more work here than any amount of writing about the mood or the feeling you are hoping to land.
Prompt
Variants worth knowing:
- Your own voice. A cloned read of your own voice suits this format unusually well.
- No music. Voice and ambience alone is a harder, colder version of the same thing.
- A series. These work as a run. Keep them in one project so the voice and rules stay identical.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real set of cinematic shorts. The voice comes first and everything else is built to it — including the footage search, which is done line by line rather than as one mood board.
- Checks the voice, music and video tools and their costs.
- Searches for a narrator matching your description in the script's language.
- Generates a test segment to hear the voice before committing.
- Tries an alternative voice when the first is not right.
- Checks the first segment's duration, then generates the remaining four in parallel.
- Checks the duration of every segment so the visuals can be built to fit.
- Searches the music library for the exact genre named in the spec and downloads it.
- Searches stock footage segment by segment — a man on the edge of a bed, a sleeping face, a clock.
- Keeps searching per line until every segment has the specific shot its words describe.
Step 8 is the difference between this format working and not. Vizard Agent searched for the literal image each line described rather than assembling a general mood, which is why the text and the picture agree with each other.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 41.0 seconds, AAC audio. Vertical, forty-one seconds, five narrated segments over sourced footage with large text on screen and the music sitting well under the voice.
If the voice is hard to hear, say so — Vizard Agent lowers the music against the stems rather than re-rendering the video, which is a fast change.
Timing was not measured separately for this kind of job. The closest measured work runs a median of 28 to 38 minutes end to end, with the middle half spread considerably wider. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.
When does this not work well?
The format is saturated and its production is easy, which puts all the weight on the script. Vizard Agent will make any script look like this, and a video that looks the part while saying nothing is the most common outcome in this genre.
- The words carry it. Generic advice in a cinematic wrapper is still generic advice.
- Music over voice is the standard failure. Name a level, and check the result with headphones.
- Stock footage is shared. The same clip appears in a hundred other shorts in this style.
- Three words on screen is a real constraint. Long sentences cannot be shown this way; the script has to be written for it.
- Advice about mental health is not neutral. This genre trades in it constantly, and what you assert is your responsibility.
How do you fix a result that came back wrong?
Name the segment or the setting that is wrong. Vizard Agent keeps every narrated segment, its measured duration, the chosen music and each downloaded clip as separate pieces, so lowering the music level or swapping a single shot leaves the read, the text and the timing of everything else exactly untouched.
- "The music is too loud." A mix change on the stems.
- "Segment three needs a different shot." One new search, same timing.
- "Slower delivery." Re-read, and the visuals re-timed to it.
How does Vizard Agent compare to doing it yourself?
By hand this is a text-to-speech tool, a stock site, a music track and an editor — five segments, five searches, five text animations, and a mix. It is a couple of hours per short, on a format that only works if you post many.
| By hand | Vizard Agent | |
|---|---|---|
| Casting the voice | Scroll a voice list | Auditioned against your spec |
| Footage per line | Search one at a time | Searched per segment, in order |
| Timing the text | Nudge each caption | Built to the measured segments |
| The mix | Re-render to adjust | Stems kept, adjust and re-mix |
Common questions
Can it write the script? Yes, from a theme. Your own writing is what separates one of these from the thousands that look identical.
Can it use my voice? Yes. Provide a clean sample and Vizard Agent clones it for the narration.
Can it work in my language? Yes. Vizard Agent searches for a narrator in that language and generates the read there.
How loud should the music be? Well under the voice. Name a level and Vizard Agent mixes to it rather than guessing.
How long should one be? Around forty seconds, which is about five segments. Longer loses the tension the format depends on.
Can I make a series with the same voice? Yes. Keep them in one project and Vizard Agent reuses the same narrator and rules.
Can it use my own footage? Yes, and it helps. Vizard Agent prefers your material and searches stock only for the lines you have not covered.
How many segments fit in forty seconds? Around five. Vizard Agent measures each narrated segment and builds the visuals to those durations rather than to even splits.
Can the text be a different colour scheme? Yes. Name the colours and what each one is for, and Vizard Agent applies that rule across every segment.