How to make a step by step animated guide to a hands on job
Name the task and ask for it step by step. Vizard Agent looks up how the job is actually done before generating a frame, builds the shots to a consistent visual style, labels each step on screen, and times the narration so the key word lands on the moment the tool makes contact.
What is the short version?
An instructional video that looks good and teaches the wrong technique is worse than no video, because someone will follow it. For this kind of video the research is not preparation for the work — it is the work, and the animation is what carries it.
- Go to Vizard Agent and name the job.
- Say you want it step by step, with the steps labelled.
- Say who it is for, so it pitches the detail correctly.
What do you need before you start?
Nothing but the task, named the way the trade names it. Vizard Agent generates the footage, the narration, the labels and the sound, so the only thing you must supply is a job specific enough that there is a right way to do it.
- The task. In the words a tradesperson would use.
- The audience. Beginner, apprentice, qualified.
- The steps, if you have them. Otherwise it researches them.
- Any safety line. Gear, warnings, what not to do.
- The platform. It sets the frame and the pace.
What do you type into Vizard Agent?
Ask for the steps explicitly. Without that word you get a nicely shot montage of somebody working, which looks like an instructional video and teaches nothing, because the viewer cannot tell where one step ends and the next begins.
Prompt
Variants worth knowing:
- "Check the technique." Research before generation.
- "Number the steps." On-screen labels, not just narration.
- "For a complete beginner." Sets how much it assumes.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real animated guide to scoring and splitting a concrete block. Steps two and three come before anything at all is drawn, which is the whole design of this job: Vizard Agent settles what the correct method is before it commits a single frame to a technique.
- Loads the explainer playbook and checks its tools.
- Researches the correct hammer-and-chisel technique.
- Checks masonry sources for the details of the method.
- Reads the generation and voiceover options.
- Generates the narration and a visual style anchor.
- Measures the narration and splits it into paced segments.
- Looks at the style anchor before generating anything else.
- Generates all ten shots against that anchor.
- Builds a contact sheet and reviews every shot together.
- Makes the step labels, a music bed and the clip audio.
- Times the key word to the moment of the split.
- Finds the exact impact frame and checks it.
- Samples the rest of the split shot and reviews the action.
- Builds the narration and sound-effect tracks.
- Designs the step chip, renders it, and moves it to the top left.
- Renders the full video and reviews frames and waveform.
- Auditions caption fonts for spacing and readability.
- Renders the final cut and watches it for sync.
Step eleven is the craft. In a manual-skill video the narration and the picture have to agree at one specific frame — the word "split" and the block actually splitting. A tenth of a second out and the viewer stops believing the video knows what it is talking about.
Step seven matters for a different reason. Ten separately generated shots will drift apart in style unless they are all anchored to the same reference, and a guide whose chisel changes shape between step three and step four is not a guide.
What does the result look like?
A numbered sequence of animated shots in one consistent look, a narration paced to the action rather than read over it, a step chip in the corner so the viewer always knows where they are, and sound effects on the contacts.
The technique in it matches what masonry sources actually describe — scoring first, working the line, then the split — rather than what an animation of a person hitting a block would plausibly look like.
When does this not work well?
Generated animation is a good teacher of sequence and a poor teacher of feel. Anything where the skill lives in pressure, sound or resistance is better filmed, and anything where being wrong is dangerous needs a qualified person to sign it off.
- Skills that live in the hands. Feel does not animate.
- Regulated or dangerous work. Have a professional check it.
- Very specific equipment. It generalises the tool.
- Jobs with local code requirements. Rules vary by country.
- Where the real thing exists. Film it; real footage wins.
How do you fix a result that came back wrong?
Correct the method, not the picture. Vizard Agent keeps the researched steps, the style anchor, the narration timings and the shot set, so a corrected technique re-generates the affected shots and re-times the narration without rebuilding the whole guide.
- "That grip is wrong." That shot re-generated to the anchor.
- "You skipped scoring." A step inserted and renumbered.
- "Too fast to follow." Re-paced with longer holds.
- "Add the safety glasses." Applied across every shot.
How does Vizard Agent compare to doing it yourself?
By hand this is either a camera, a block and an afternoon, or an animator who does not know the trade working from a script written by someone who does. Both work. Both are a week of arranging things before anything gets made.
| By hand | Vizard Agent | |
|---|---|---|
| The technique | You know it, or you look it up | Researched from trade sources first |
| Consistency across shots | One shoot, one setup | Every shot anchored to one reference |
| Step labels | Added in post | Designed and placed, then checked |
| Narration sync | Read over the picture | Key word timed to the impact frame |
| Changing a step | Reshoot | Re-generate that shot |
Common questions
Does it actually look things up? Yes. Vizard Agent researches the method before generating anything.
Can I give it my own steps? Yes, and it will follow them. Vizard Agent researches only what you leave open.
Will the tool look right? It is anchored to one reference. Vizard Agent keeps it consistent across shots.
Can it use my footage instead? Yes. Vizard Agent will cut real footage into the same numbered structure.
How many steps can it do? As many as the job has. Ten shots is a comfortable short-form length.
Does it add safety warnings? If you ask, or if the sources call for it. Tell Vizard Agent what must appear.
Can it do it in another language? Yes. Vizard Agent narrates and labels in the language you name.
Will it show the mistake to avoid? Yes, if you ask for a "do not do this" beat.
Can it match my training style? Give it an existing video and Vizard Agent works to that look.
Is the narration timed to the action? Yes. That is the part Vizard Agent spends real effort on.
Can it be longer than a short? Yes. Vizard Agent paces the same structure across a longer runtime.
Does it check its own render? Yes. Vizard Agent reviews frames and the waveform before finalising.
Can I get the script separately? Yes. Ask and Vizard Agent delivers the narration as text.
Should a professional review it? For anything regulated or dangerous, yes. Research is not qualification.