Vizard Agent

How to make a play-along brain test video

Last updated 2026-08-30 · 7 min read

Write the test and tell Vizard Agent the viewer plays along out loud. Vizard Agent casts an energetic voice, pulls word-level timings from the narration, generates the visuals and sound effects together, and times every reveal so the answer lands after the viewer has already committed to theirs.

What is the short version?

A test video works entirely because the viewer answers before you do. Everything about the edit serves that one mechanism: the prompt appears, there is a beat where they say their answer out loud, and only after that does the reveal arrive on screen.

  1. Go to Vizard Agent and write the test, including the reveal.
  2. Say the viewer should be playing along out loud.
  3. Say the pace — quick-fire, or with room to think.

What do you need before you start?

The test and its trick. Vizard Agent builds the visuals, the voice and the sound, so what it needs from you is the mechanism — what makes it harder than it looks — because that is what the timing has to protect.

What do you type into Vizard Agent?

Write it as you would say it out loud. Vizard Agent builds the video from the narration and the timings inside it, so writing the whole thing in a speaking voice — including the "ready?" — gets you a piece that already has the rhythm the format needs.

Prompt

Make a play-along brain test video. Script: "[Guess the colour. Your brain thinks this is easy until the colours start lying. Say the shade, not the word. Ready?]" then [each round]. Energetic voice, bold captions, quick pace.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real play-along test. The step that sets the rhythm is the eighth: it pulled word-level timings from the narration, so the reveal can be placed relative to the exact word that sets each round up.

  1. Checks the motion graphics and explainer template guidance.
  2. Checks the available capabilities and their pricing.
  3. Checks the voiceover options and the caption, motion-graphics and music tools.
  4. Searches for an energetic voice suited to a game.
  5. Generates the voiceover and searches for background music.
  6. Downloads the music and transcribes the voiceover for word timestamps.
  7. Inspects the timing segments and then the individual word timings.
  8. Checks the sound-effect and image generation options.
  9. Generates the visual scenes and the sound effects at the same time, reviews the generated images, checks the installed bold typefaces and builds styled animated captions.

Step 9's parallel generation matters at this length. The scenes and the sound do not depend on each other, so building them together rather than in sequence is most of why a sixty-second game does not take an afternoon.

What does the result look like?

A short vertical video that talks directly to the viewer: an instruction, a warning that it is harder than it looks, then rounds arriving quickly with bold captions, a sound on each one, and the reveal landing just after the point where somebody watching would have answered.

The pacing is the whole design. Half a second too early and the video answers for them; a beat too late and the momentum goes, which is why the reveals are placed against word timings rather than a fixed interval.

When does this not work well?

A play-along format asks something of the viewer rather than just showing them something, and it fails whenever that asking does not land. These are the situations where Vizard Agent builds the video correctly and the format still does not work.

How do you fix a result that came back wrong?

Say which round is wrong and what should happen instead. Vizard Agent keeps the word timings, all the generated scenes, the sound effects and the caption styling, so the pacing or a single round can change without a rebuild.

How does Vizard Agent compare to doing it yourself?

By hand this means building each round as its own graphic, recording narration, and then nudging every reveal until the gap feels right — which is a judgement you make twenty times and get slightly wrong in at least three of them.

By hand Vizard Agent
Timing the reveals Nudge until it feels right Placed against word timings
The rounds Build each graphic Generated with the sound in parallel
The voice Read it yourself Cast for energy, then timed
Captions Whatever font is loaded A bold typeface located and styled

Common questions

What makes a play-along video work? The gap before the reveal. Vizard Agent times that gap so the viewer answers first.

Does it need sound? It helps, though many watch silently, so Vizard Agent makes the captions carry it too.

How many rounds should it have? Enough to establish the trick and then punish it. Vizard Agent finds four or five works.

Can Vizard Agent generate the visuals? Yes, alongside the sound effects, in the same pass.

How long should it be? Thirty to sixty seconds. Vizard Agent keeps it tight because the tension goes if it runs on.

Why time the reveals to the words rather than to seconds? Because the viewer answers in response to a specific word, not at a fixed moment, and a reveal placed on the clock lands early on one round and late on the next.

Can I use my own test? Yes. Vizard Agent builds whatever mechanism you write, provided the trick is clear.

Does it work for children? Yes. Tell Vizard Agent the age and it uses a simpler test at a slower pace.

Is this an actual assessment of anything? No, and the video should not imply otherwise. It is a game.

Should the video tell people the answer at all? Yes, and immediately after the gap. Vizard Agent places the reveal where it confirms or contradicts what the viewer just said, which is the moment the whole format exists to produce.

Why generate the visuals and the sound together? Because neither depends on the other. Vizard Agent builds them in parallel and aligns them afterwards against the word timings, which is most of why a sixty-second game does not take an afternoon.

Does Vizard Agent check the result? Yes. It reviews the generated scenes and the caption timing before delivery.