How to turn a piece of writing into a calm narrated video
Paste your text into Vizard Agent and say how it should feel. Vizard Agent casts and records the narration, generates the imagery, times every shot to the narration's own word timings, builds subtitles and a closing card, then checks the finished cut's frames and audio levels before re-rendering anything that falls short.
What is the short version?
A piece of writing you care about becomes a video by being read aloud well and given pictures that stay out of its way. The failure mode is imagery that is too busy and a voice that is too performed.
- Go to Vizard Agent and paste the text.
- Say the voice — calm, warm, unhurried.
- Say the feel of the imagery and where it is going.
What do you need before you start?
The text and a register. Vizard Agent handles the voice, the pictures and the timing, so the decisions that stay with you are the words themselves and how they should be read — which for this format is nearly the whole job.
- The text, as written. Line breaks and pacing marks are read as pacing.
- Any timings you want. If you have marked when each line should land, Vizard Agent will follow them.
- The voice. Calm and low is the convention here. Say the age and register.
- The imagery. Generated, abstract, or a person on screen. This changes the video completely.
- The length. Around forty seconds is what a short piece of writing supports.
What do you type into Vizard Agent?
Paste the words and describe the tone in one line. Vizard Agent will not rewrite your text unless you ask, so the whole brief is the writing plus a sentence about the voice and how still or flowing the pictures should be.
Prompt
Variants worth knowing:
- A person on screen. Ask for a character and Vizard Agent builds a reference sheet so they stay the same throughout.
- No subtitles. Some pieces read better without words on screen; ask if you want them left off.
- Your own voice. Provide a sample and Vizard Agent clones it rather than casting a narrator.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real narrated piece. The end of the run is the part worth reading: it looked at the finished video, measured its sound, and decided it was not good enough yet.
- Loads the mood-montage playbook and reads the generation options.
- Generates the narration, searches for music and produces the first batch of shots in one pass.
- Builds a character reference sheet and measures the narration's timings.
- Reads those word timings and assembles a contact sheet of the generated shots.
- Looks at the shots, then generates the ones featuring the person and reviews them too.
- Listens to both music candidates and browses the available fonts.
- Checks the word timings against the clip lengths and assembles the narration track and the picture cut.
- Generates the subtitles and the closing title card, then renders.
- Samples frames from the finished cut, checks the audio waveform, measures the levels, and re-renders with better subtitle contrast and a balanced mix.
Step 9 is the difference between a first render and a delivery. Vizard Agent measured its own output — the picture and the sound — and went back rather than treating the first complete file as the finished one.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 24fps, 37.0 seconds, AAC audio. Vertical, thirty-seven seconds, narrated over generated imagery with subtitles and a closing card, and the music ducked under the voice.
Vizard Agent watched the finished video back after the final render as a last check rather than uploading straight from the encoder.
This category was not separately measured, so treat the timing as a range rather than a promise: comparable work runs a median of 28 to 38 minutes end to end. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.
When does this not work well?
The text carries this format completely, and everything else in the video is in service of it. Vizard Agent gives the writing a well-cast voice, quiet imagery and clean subtitles, and none of that treatment makes a piece of writing any better than it was when you pasted it in.
- Weak writing shows more, not less. A calm read and slow pictures leave nowhere for a thin line to hide.
- Busy imagery fights the words. If the pictures are interesting, the text stops being heard.
- A generated person can distract. Sometimes abstract imagery serves a personal piece better than a face.
- Someone else's writing is theirs. Setting a published poem or lyric to video is a permission question.
- Subtitles change the register. They help reach and they make a quiet piece feel like a social post; consider leaving them off.
- Forty seconds is a short piece. A longer text read at this pace runs past what a feed will hold, and splitting it works better than speeding it up.
How do you fix a result that came back wrong?
Say which element. Vizard Agent keeps the narration, its word timings, every generated shot, both music candidates and the subtitle file separately, so a slower read or a swapped shot is a change to one piece rather than a rebuild of the whole cut.
- "The voice is too performed." A flatter, calmer read, everything re-timed.
- "That image is too literal." One shot regenerated at the same timing.
- "Subtitles are hard to read." Restyled for contrast over the same render.
How does Vizard Agent compare to doing it yourself?
By hand this is a text-to-speech pass, a search for imagery that is not trite, a subtitle file typed and timed, and a mix where the music sits over the voice because you stopped checking after the first export.
| By hand | Vizard Agent | |
|---|---|---|
| The read | Generic text-to-speech | Cast to a described register |
| Imagery | Stock that nearly fits | Generated to the piece |
| Timing | Nudge each shot | From the narration's word timings |
| The final check | Play it once | Frames sampled, levels measured, re-rendered |
Common questions
Will it change my words? No. Vizard Agent reads the text exactly as written unless you ask for an edit, including the line breaks, which it treats as pacing.
Can it put a closing card at the end? Yes. Vizard Agent builds one and checks its contrast against the imagery behind it.
Can it use my own voice? Yes. Give a clean sample and Vizard Agent clones it for the narration.
Can the same person appear in every shot? Yes. Vizard Agent builds a character reference sheet before generating anything, which is what holds the same person across every shot in the video.
Do I have to have subtitles? No. Say so and Vizard Agent leaves them off, which suits a quieter piece.
How long should it be? Around forty seconds for a short piece. Longer texts are better split into several videos than read faster to fit one.
Can I supply timings for each line? Yes. Vizard Agent follows the timings you mark rather than deciding the pacing itself.
Can it add music? Yes, and it ducks the track under the voice rather than sitting it on top, which matters more in this format than any other.