Vizard Agent

How to make an illustrated story video for children

Last updated 2026-08-16 · 5 min read

Tell Vizard Agent which story to tell and what the narrator should sound like. Vizard Agent casts a storyteller voice, records the narration first, measures its word-level timing, and then generates each illustration so the pictures arrive on the words they belong to rather than on a fixed schedule.

What is the short version?

A children's story video is narration with pictures attached, and the order in which you make those two things decides whether it works. Build the pictures first and the reading has to be squeezed to fit them. Record the voice first and the pictures can land exactly where the story turns.

  1. Go to Vizard Agent and name the story or paste your text.
  2. Say what the storyteller should sound like.
  3. Say the language, the length and the age it is for.

What do you need before you start?

A story and a voice. Vizard Agent knows the classic fables and will write from your own text just as readily, so the decision that shapes the video most is the narrator — a warm grandmother and a brisk teacher produce two very different films from identical words.

What do you type into Vizard Agent?

Name the story and the voice. Everything else — the illustration style, the pacing, where the pictures change — Vizard Agent derives from the narration once it exists, which is a better order than deciding the look before you know how the reading sounds.

Prompt

Make a storytelling video for children based on [the fable], narrated by [a warm, enthusiastic storyteller]. In [English], around [1] minute.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real children's story. The sequence itself is the lesson worth taking away: the voice comes first, its timing is measured rather than estimated, and only then are the pictures built to land against those measurements.

  1. Lists the available capabilities to see what it has to work with.
  2. Loads its film-making guidance rather than improvising a structure.
  3. Checks the image generation and voiceover options before choosing anything.
  4. Searches for a warm, enthusiastic female storyteller voice suitable for a children's story.
  5. Generates the narration in the cosy storytelling voice it settled on.
  6. Measures the exact duration of the generated narration file.
  7. Reads the transcription tool's options for word-level timing.
  8. Transcribes its own narration to get exact word-level timestamps.
  9. Generates the first scene illustration — the thirsty crow in bright daylight — against those timings.

Step 8 is the part worth stealing. Vizard Agent transcribed the voice it had just generated, purely to find out precisely when each word lands, so the pictures could be cut to the story rather than to a stopwatch.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1920x1080, H.264, 25fps, 50.08 seconds, AAC audio. Wide rather than vertical, which suits a story watched on a television or a tablet rather than scrolled past in a feed.

Ask for a vertical version in the same conversation and Vizard Agent re-lays the illustrations for the new shape, using the narration and the timings it already has.

It is not instant, and this category was not separately measured. Comparable work runs a median of 28 to 38 minutes end to end. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.

When does this not work well?

Illustrated stories are generous to generation, since nobody watching expects photographic accuracy from a drawing. What still needs watching is consistency between scenes and the content of the story itself, and neither is something Vizard Agent will flag on your behalf.

How do you fix a result that came back wrong?

Name the scene or the line that is wrong. Vizard Agent keeps the narration, the word-level timings it measured and each illustration as separate pieces, so a single picture can be regenerated without re-recording any of the story, and every other picture stays exactly where it was placed.

How does Vizard Agent compare to doing it yourself?

The manual version is a text-to-speech tool, a stack of illustrations from somewhere, and an editor to time them together. The timing is the part that takes the evening: every picture has to change on the right word, and every change to the narration breaks all of them.

By hand Vizard Agent
Narration A TTS preset, exported once Cast, generated, then measured
Illustrations Bought, drawn or generated separately Built against the narration
Timing pictures to words By ear, one at a time From word-level timestamps
Changing a line Re-time everything after it Re-read that line, re-time

Common questions

Does it know the classic fables? Yes. Name one and Vizard Agent writes and tells it, or paste your own text if you want a specific version.

Can it tell the story in my voice? Yes. Upload a clean sample of ten to ninety seconds and the story is told in it.

Will the characters look the same throughout? Say so and Vizard Agent builds a reference to hold them. Without that instruction, scenes are generated independently.

Can it put the words on screen too? Yes. Ask for captions and Vizard Agent uses the same word-level timings it measured for the pictures, so the text and the reading stay together.

Can I get a series? Yes. Keep them in one project and Vizard Agent matches the voice and the illustration style across every story.