Vizard Agent

How to turn a single diagram into a moving explainer

Last updated 2026-08-25 · 7 min read

Upload the diagram and talk through what each panel means. Vizard Agent reads the image, maps its coordinates, builds a focus still for every panel, times badges to the word-level narration, and cuts an explainer that moves across the diagram in the order you explained it.

What is the short version?

A three-panel diagram already contains the argument. What it lacks is time — the viewer sees all three panels at once and reads none of them, so the explainer's job is to take them one at a time in the right order.

  1. Go to Vizard Agent and upload the diagram.
  2. Explain each panel in your own words, left to right.
  3. Point at your earlier videos so the style matches.

What do you need before you start?

The image and the explanation. Vizard Agent reads the diagram itself and maps where every panel sits, so describing the layout is unnecessary. What it cannot infer is meaning, so you do need to say what each panel represents and why it is there.

What do you type into Vizard Agent?

Talk it through panel by panel rather than writing a polished script. Vizard Agent turns your spoken explanation into narration and uses the order you said things in as the order the camera moves across the diagram, so the sequence follows your reasoning.

Prompt

Turn this [3]-panel graphic into a [50] second explainer. Left panel is [the problem], middle is [the change], right is [the outcome]. Match the style of my earlier [product] videos.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real three-panel explainer. The step that carries the whole piece is twenty-fourth: it builds a separate focus still for each panel, so the frame can actually move rather than sitting on the full image.

  1. Looks at the three-panel screenshot you attached.
  2. Checks your past projects for the brand and style and loads the software-ad playbook.
  3. Downloads the graphic and your previous videos on the same product.
  4. Extracts frames from those earlier ads and reviews them — openings, mid-shots, endings.
  5. Maps the panel coordinates and reads the coordinate grid over the graphic.
  6. Listens to the previous voiceover and music to match the house sound.
  7. Generates the narration, downloads the music, and builds left, middle and right focus stills.
  8. Checks each focus still by eye, then remakes them with timing badges against the word-level narration timings.
  9. Generates business-style captions and a whoosh, checks the badge designs, and cuts the explainer to the word timings.

Step 7 is the difference between an explainer and a still image with a voice over it. Three separate crops of the same graphic give the edit somewhere to go, and the badges land on the word that names the number.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1920x1080, H.264, 30fps, 50.6 seconds, AAC audio. Widescreen, fifty seconds, one static graphic turned into a moving explainer with three focus positions, timed badges, business captions and music matched to the earlier videos in the same series.

Fifty seconds is about right for a three-panel argument. Each panel gets fifteen seconds, which is long enough to read and short enough that the viewer does not leave before the third one.

When does this not work well?

A diagram has to be legible before Vizard Agent can explain it, and cropping magnifies whatever is already there. Most of the failures on this format start with the source image rather than with anything the edit did wrong.

How do you fix a result that came back wrong?

Name the panel that is wrong. Vizard Agent keeps the coordinate map, the three focus stills, the badge designs and the word-level timings, so a single crop can be remade or one badge re-timed without rebuilding the whole explainer.

How does Vizard Agent compare to doing it yourself?

By hand this means cropping the graphic three times in an image editor, recording narration, then keyframing each crop so the moves land on the right words. The badge timing is the fiddly part, and it is the part that most explainers get slightly wrong.

By hand Vizard Agent
Panel crops Cut them in an image editor Focus stills from a read coordinate map
Narration timing Nudge keyframes by ear Cut to word-level timings
Badges Place and re-place Timed to the word that names the number
House style Rebuild from memory Read from your earlier videos

Common questions

How many panels can this handle? Three fits comfortably in fifty seconds. Tell Vizard Agent to allow more time for more panels rather than cutting faster.

Do I need to write a script? No. Talk through what each panel means and Vizard Agent builds the narration out of it.

Will it match my other videos? Yes, if you point at them. Vizard Agent reads their frames, voice and music before starting.

What resolution does the diagram need? High enough to survive cropping. The focus stills Vizard Agent builds magnify whatever detail is in the source.

Can numbers appear as badges? Yes. Vizard Agent times them to the word that says them rather than dropping them at a fixed second.

Why crop the image at all? Because a static full-frame graphic gives the viewer everything at once and they read none of it. Three focus positions give the explanation somewhere to go and let each panel land on its own.

Can I get a vertical cut? Yes. Vizard Agent re-frames the focus stills for the narrower shape.

Does it work for flowcharts and architecture diagrams? Yes, provided the source is legible. Vizard Agent uses the same method for any panelled graphic.

Can the captions be subdued? Yes. Ask for a business style and Vizard Agent keeps them restrained.

Can Vizard Agent use my existing narration instead? Yes. Upload the audio and Vizard Agent pulls word-level timings from it, then times the panel moves and the badges to what you actually said rather than generating a new voice for the piece.

Does Vizard Agent check the finished explainer? Yes. It reviews the crops and the badge timing before delivery.