Vizard Agent

How to make a comedy micro-drama where only one character speaks

Last updated 2026-08-20 · 6 min read

Give Vizard Agent the script and say which character is the only audible one. Vizard Agent auditions voices for that character, generates every other character's reactions as silent performance, builds the on-screen interface as clean vector graphics, and cuts the timing to the spoken lines.

What is the short version?

A one-voice micro-drama is a real format with a real constraint: one character talks and everyone else reacts. That constraint is what makes it funny and it means the silent characters need generated expressions precise enough to carry a beat without a line.

  1. Go to Vizard Agent and give it the script and the beats.
  2. Say which character is the only one who speaks.
  3. Give a reference image for the look and the layout.

What do you need before you start?

The script and a reference. Vizard Agent generates the characters and the interface itself, and the two things it cannot infer are the comic timing you have in mind and the visual world the episode lives in, which a single reference image conveys faster than a description.

What do you type into Vizard Agent?

Write the rule as a rule. Vizard Agent treats "she is the only audible speaking character" as a hard constraint on the whole build, which changes how the other characters are generated and how the beats are timed around them.

Prompt

Make a [45-60] second comedy micro-drama from this script. [Grandma] is the only audible speaking character — everyone else reacts silently. Match the layout in the reference image.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real micro-drama episode. Two things stand out: it auditioned voices before generating a single line, and when the generated interface came back wrong it rebuilt it as vectors instead.

  1. Opens the reference image and inspects the layout dimensions.
  2. Slices the reference into panels and crops the character faces out of them.
  3. Searches for candidate voices, tests them and listens back before choosing.
  4. Generates the speaking character's lines and measures the total duration.
  5. Generates the silent character's performance — the guilty look, the gesture, the suppressed laugh — and reviews each batch as tiles.
  6. Generates clean character portraits and key close-up frames and inspects them.
  7. Tests the split-frame layout logic, generates a full interface frame and looks at it.
  8. Rebuilds the interface as clean vector graphics after the generated version did not hold up, and re-checks it.
  9. Measures the exact voiceover timings, sources comedy sound effects for the beats, measures the loudness of each audio stem and renders the episode to the spoken timing.

Step 8 is a judgement most pipelines cannot make. A generated screenshot of an app looks approximately right and reads as fake; drawing the interface as vectors is slower to decide on and correct at every resolution.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 60.98 seconds, AAC audio. Vertical, just over a minute, one audible voice, generated reaction shots for the silent character, and a vector interface layered over the whole episode.

Running to sixty-one seconds against a forty-five to sixty second brief is the voice setting the length. Vizard Agent times the cut to the measured narration rather than trimming a punchline to hit a round number.

When does this not work well?

Generated characters and comic timing are each difficult on their own, and this format needs both of them working at once. Vizard Agent reviews every image it generates before using it, and some things stay outside what it can reasonably produce.

How do you fix a result that came back wrong?

Name the beat that is off. Vizard Agent keeps the voice takes, every generated character frame, the vector interface graphics and the measured word timings, so re-cutting a joke or swapping one reaction is a re-render against material Vizard Agent already produced.

How does Vizard Agent compare to doing it yourself?

By hand this is casting a voice, generating dozens of character images that do not match each other, mocking up an interface in a design tool, and then cutting the whole thing to timing you have to measure yourself.

By hand Vizard Agent
Casting the voice Pick one and commit Candidates tested and listened to
Silent reactions Generate and hope Generated in batches, reviewed as tiles
The interface Screenshot or mock-up Drawn as vectors after the first attempt failed
Comic timing Nudge clips by hand Cut to measured word timings

Common questions

Why only one speaking character? It is a format convention, and it makes the silent reactions do the comedy.

Can it keep the character across episodes? Yes, from the same reference. Long series still drift over time.

Does it add sound effects? Yes, on the comedy beats where a second voice would otherwise be carrying the moment. Vizard Agent measures the loudness of each audio stem so the stings sit under the voice rather than over it.

Can it build a fake app interface? It builds a stylised one as vectors. A replica of a real app's branding is not something to ship.

How long should an episode be? Forty-five to sixty seconds. Vizard Agent lets the voice set the exact length.

Do I need a reference image? It helps a great deal. Vizard Agent slices it into panels and crops the characters and the layout out of it.

How does it choose the voice? Vizard Agent searches for candidates, generates test lines and listens to them before recording the script.

Does it measure the timing? Yes. Vizard Agent measures the exact voiceover timings and cuts the picture to them rather than the other way round.

Can I ask for a series? Yes. Vizard Agent generates clean character portraits specifically so that later episodes have something consistent to work from.

How does it handle the on-screen comments? As part of the vector interface, timed to the beats, so the comments land as jokes rather than sitting there as decoration.