How to make a comedy micro-drama where only one character speaks
Give Vizard Agent the script and say which character is the only audible one. Vizard Agent auditions voices for that character, generates every other character's reactions as silent performance, builds the on-screen interface as clean vector graphics, and cuts the timing to the spoken lines.
What is the short version?
A one-voice micro-drama is a real format with a real constraint: one character talks and everyone else reacts. That constraint is what makes it funny and it means the silent characters need generated expressions precise enough to carry a beat without a line.
- Go to Vizard Agent and give it the script and the beats.
- Say which character is the only one who speaks.
- Give a reference image for the look and the layout.
What do you need before you start?
The script and a reference. Vizard Agent generates the characters and the interface itself, and the two things it cannot infer are the comic timing you have in mind and the visual world the episode lives in, which a single reference image conveys faster than a description.
- The script. Lines for the speaker, beats for everyone else.
- The speaking rule. Who is audible and who is not.
- A reference image. The characters and the on-screen layout.
- The length. Forty-five to sixty seconds per episode.
- The tone. Comedy, drama, chaos, or all three.
What do you type into Vizard Agent?
Write the rule as a rule. Vizard Agent treats "she is the only audible speaking character" as a hard constraint on the whole build, which changes how the other characters are generated and how the beats are timed around them.
Prompt
Variants worth knowing:
- A recurring character. Say so and later episodes stay consistent.
- On-screen chat. Comments scrolling past are half the joke in this format.
- Sound effects on the beats. Stings do the work a second voice would.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real micro-drama episode. Two things stand out: it auditioned voices before generating a single line, and when the generated interface came back wrong it rebuilt it as vectors instead.
- Opens the reference image and inspects the layout dimensions.
- Slices the reference into panels and crops the character faces out of them.
- Searches for candidate voices, tests them and listens back before choosing.
- Generates the speaking character's lines and measures the total duration.
- Generates the silent character's performance — the guilty look, the gesture, the suppressed laugh — and reviews each batch as tiles.
- Generates clean character portraits and key close-up frames and inspects them.
- Tests the split-frame layout logic, generates a full interface frame and looks at it.
- Rebuilds the interface as clean vector graphics after the generated version did not hold up, and re-checks it.
- Measures the exact voiceover timings, sources comedy sound effects for the beats, measures the loudness of each audio stem and renders the episode to the spoken timing.
Step 8 is a judgement most pipelines cannot make. A generated screenshot of an app looks approximately right and reads as fake; drawing the interface as vectors is slower to decide on and correct at every resolution.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 60.98 seconds, AAC audio. Vertical, just over a minute, one audible voice, generated reaction shots for the silent character, and a vector interface layered over the whole episode.
Running to sixty-one seconds against a forty-five to sixty second brief is the voice setting the length. Vizard Agent times the cut to the measured narration rather than trimming a punchline to hit a round number.
When does this not work well?
Generated characters and comic timing are each difficult on their own, and this format needs both of them working at once. Vizard Agent reviews every image it generates before using it, and some things stay outside what it can reasonably produce.
- Character consistency drifts. Vizard Agent generates clean reference portraits to hold a look, and long series still drift.
- Comedy is specific. A beat that works in your head needs to be written as a beat, not implied.
- Real people are off limits. Generating a likeness of a real person is not something to build a series on.
- Platform interfaces are trademarks. A stylised interface is fine; a replica of a real app's branding is not.
- One voice carries everything. If the audition does not land, the episode does not either.
How do you fix a result that came back wrong?
Name the beat that is off. Vizard Agent keeps the voice takes, every generated character frame, the vector interface graphics and the measured word timings, so re-cutting a joke or swapping one reaction is a re-render against material Vizard Agent already produced.
- "The pause before the punchline is too short." Re-timed against the word timings.
- "His reaction is wrong there." Regenerated from the same character reference.
- "The voice is too young." Another candidate from the audition, already tested.
How does Vizard Agent compare to doing it yourself?
By hand this is casting a voice, generating dozens of character images that do not match each other, mocking up an interface in a design tool, and then cutting the whole thing to timing you have to measure yourself.
| By hand | Vizard Agent | |
|---|---|---|
| Casting the voice | Pick one and commit | Candidates tested and listened to |
| Silent reactions | Generate and hope | Generated in batches, reviewed as tiles |
| The interface | Screenshot or mock-up | Drawn as vectors after the first attempt failed |
| Comic timing | Nudge clips by hand | Cut to measured word timings |
Common questions
Why only one speaking character? It is a format convention, and it makes the silent reactions do the comedy.
Can it keep the character across episodes? Yes, from the same reference. Long series still drift over time.
Does it add sound effects? Yes, on the comedy beats where a second voice would otherwise be carrying the moment. Vizard Agent measures the loudness of each audio stem so the stings sit under the voice rather than over it.
Can it build a fake app interface? It builds a stylised one as vectors. A replica of a real app's branding is not something to ship.
How long should an episode be? Forty-five to sixty seconds. Vizard Agent lets the voice set the exact length.
Do I need a reference image? It helps a great deal. Vizard Agent slices it into panels and crops the characters and the layout out of it.
How does it choose the voice? Vizard Agent searches for candidates, generates test lines and listens to them before recording the script.
Does it measure the timing? Yes. Vizard Agent measures the exact voiceover timings and cuts the picture to them rather than the other way round.
Can I ask for a series? Yes. Vizard Agent generates clean character portraits specifically so that later episodes have something consistent to work from.
How does it handle the on-screen comments? As part of the vector interface, timed to the beats, so the comments land as jokes rather than sitting there as decoration.