Vizard Agent

How to turn handwritten notes into a lesson video in your own voice

Last updated 2026-09-01 · 8 min read

Photograph your notes and record about thirty seconds of yourself speaking. Vizard Agent clones the voice for the full narration, detects where every line of writing sits on each page, and fades each item in at the exact moment the narration reaches it.

What is the short version?

A teacher's board notes already contain the lesson. What they lack is the sequence — the fact that the third line only makes sense after the second. Revealing them one at a time, timed to the explanation, is the whole difference between a photograph and a lesson.

  1. Go to Vizard Agent and upload photographs of the notes.
  2. Record about thirty seconds of your own voice.
  3. Say what each page teaches and in what order.

What do you need before you start?

Clear photographs and a short voice sample. Vizard Agent clones the voice from that sample and writes the rest of the narration in it, so half a minute of ordinary speaking is enough — it does not need a studio recording.

What do you type into Vizard Agent?

Say what is on the pages and how it should unfold. Vizard Agent works out the layout detection and the timing itself, so the useful part of the brief is the teaching order and any word a cloned voice is likely to get wrong.

Prompt

Make a vertical [9:16] teaching short from these [3] photos of my board notes on [autosomal and sex-linked allele combinations]. Use the attached [33] seconds of my own voice — clone it for the whole narration, about [95] seconds. Hand-drawn notebook styling. Reveal each item on the page as I reach it, not all at once. Check the pronunciation of [the technical terms] before recording everything. Captions in a style that suits handwriting.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real genetics lesson built from three photographs of board notes and thirty-three seconds of the teacher's own voice. Step seven is the one that makes the whole staged reveal possible, and step thirteen is where it pays off.

  1. Looks at the uploaded audio and photographs and checks their sizes and duration.
  2. Transcribes the voice sample and views each of the three images.
  3. Confirms the voice-cloning parameters and tests the clone's pronunciation.
  4. Listens back to the test before committing to the full narration.
  5. Generates every narration segment in parallel in the cloned voice.
  6. Finds a handwriting typeface that matches the source material.
  7. Automatically detects the position of every line of text in the photographs.
  8. Assembles the complete narration track and generates the animation cards, music and effects.
  9. Reads the word-level timings to locate every animation trigger point.
  10. Verifies a specific term's pronunciation by listening to that passage.
  11. Checks the transparent animation layers actually contain content and previews them.
  12. Measures the real bounding box of each overlay and adjusts the offsets and scale.
  13. Splits one page's twenty genotype combinations into twenty separate layers, each fading in at its own timestamp.
  14. Moves the counter label inward so the platform's interface cannot cover it.
  15. Mixes the narration up to its target loudness with the music well beneath it.

Step thirteen is what a good teacher does on a board and a slideshow never does. Twenty combinations appearing together is a wall of symbols; twenty appearing one at a time as they are named is a lesson someone can follow.

What does the result look like?

The lesson this page is written from ran about 95 seconds vertical, narrated end to end in the teacher's own cloned voice from a 33-second sample, over hand-drawn notebook styling built from three photographs of her board notes.

Each item appears as it is spoken. The overlay layers were measured and repositioned so none of them covered the handwriting underneath, and the counter on the right was moved in from the edge so the platform's own buttons could not sit on top of it.

When does this not work well?

The finished lesson can only ever be as clear as the notes it was built from, and cloned voices have one specific and predictable weakness: subject vocabulary. Vizard Agent checks the pronunciation of the technical terms up front rather than discovering the problem in the final render.

How do you fix a result that came back wrong?

Say which item on the page is wrong, or which word came out mispronounced. Vizard Agent keeps the detected line positions, every layer's measured bounding box and the full word timings, so a re-ordered reveal or a corrected pronunciation is a re-render of that one page.

How does Vizard Agent compare to doing it yourself?

By hand this becomes a slide deck carrying twenty separate build steps, each one hand-timed against a recording, where changing any single item means re-timing everything after it. Vizard Agent triggers each reveal directly from the narration's own word timings instead.

By hand Vizard Agent
The narration Recorded in full, repeatedly Cloned from thirty seconds
Reveal timing Hand-keyed per build Triggered at word timestamps
Layout Rebuilt as slides Detected from your photographs
Overlay collisions Nudged by eye Bounding boxes measured

Common questions

How much of my voice does it need? About thirty seconds. Vizard Agent clones from a short ordinary sample.

Will it sound like me? Closely. Vizard Agent tests the clone and listens back before recording everything.

What about technical terms? Flag them. Vizard Agent checks their pronunciation specifically and re-records if wrong.

Do the photos need to be good? Sharp and even. Vizard Agent detects each line of text from the image.

Can it reveal items one at a time? Yes, and this is the point. Vizard Agent triggers each from the narration's word timings.

How many pages can it handle? As many as you upload. Vizard Agent builds each as its own section.

Will overlays cover my writing? It measures to prevent that. Vizard Agent checks every layer's real bounding box.

Does it add captions? Yes, in a style that suits handwriting. Vizard Agent keeps them clear of the notes.

Can I use printed notes instead? Yes. Vizard Agent detects lines of print as readily as handwriting.

What about the music? Well underneath. Vizard Agent mixes it far below the narration so the lesson is audible.

Is cloning my own voice a problem? Not your own, no. Do not clone a colleague's without asking — a teacher's voice is recognisable to every one of their students.

Can I make a series this way? Yes. Vizard Agent reuses the cloned voice and the styling for every subsequent lesson.

Does it check the result? Yes. Vizard Agent previews each page's final state and verifies the layers before rendering.