Vizard Agent

How to make one unbroken talking-avatar video with no cuts

Last updated 2026-09-01 · 8 min read

Upload the portrait and say it must be one continuous performance. Vizard Agent auditions voices by listening to them rather than trusting their labels, generates the performance in parts to get around the generator's length limit, and joins them so the result plays as a single unbroken take.

What is the short version?

A story told straight to camera by one person, without a single cut, is the most convincing thing short-form video does — and it is the hardest to generate, because every cut is a chance for the face and the voice to drift.

  1. Go to Vizard Agent and upload the portrait.
  2. Say the rules: one person, one voice, one room, no cuts.
  3. Give the story written the way someone would actually say it.

What do you need before you start?

A portrait and a script that sounds like speech. Vizard Agent preserves the face you upload rather than reinterpreting it, so the image is the character, and the script matters more here than in any other format because nothing else is carrying the video.

What do you type into Vizard Agent?

Write the constraints as a list of things not to do. Vizard Agent will otherwise reach for cutaways and B-roll because that is what usually improves a video, and here every one of them destroys the effect you are after.

Prompt

Make ONE continuous [9:16] Short using the uploaded [elderly woman] as the ONLY character. It must feel like a real person telling this story to the viewer, not an avatar reading a script — not a narrator, not a motivational speaker, not an announcer. Keep her exact face, age, hair, clothing and proportions. Do not redesign or replace her. ONE person, ONE voice, ONE location, ONE camera, ONE continuous performance. No scene changes, no B-roll, no cutaways, no second character, no voice switching. She stays seated in the same room, looking at the viewer. Story: [...]

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real continuous-performance Short. Forty of the seventy-nine steps are casting the voice — searching, generating tests and having each one listened to — because a voice labelled "elderly" frequently is not.

  1. Inspects the uploaded portrait to understand the character.
  2. Checks the avatar and voiceover tools and their options.
  3. Searches the voice library for elderly female voices, by accent and then without.
  4. Tests candidate voices on a short line and listens to each result.
  5. Hunts for a clone sample across stock libraries and public recordings.
  6. Has the generated audio analysed for age, tone and rasp rather than trusting the label.
  7. Settles on a voice that passes the listening test and generates the full performance.
  8. Attempts the avatar in one pass and hits the generator's length limit.
  9. Splits the voiceover into two parts and generates both in parallel.
  10. Crops the portrait to an exact 9:16 frame so the avatar generates cleanly.
  11. Generates both halves from the same optimised image and concatenates them seamlessly.
  12. Generates a subtle atmospheric music bed for underneath.
  13. Transcribes the finished performance and reads the word-level timings.
  14. Generates word-by-word captions and mixes them with the music.
  15. Extracts frames across the result and reviews the composition and captions.

Step six is the one worth copying. Vizard Agent had the generated audio listened to and assessed for age and rasp, because voice libraries label by intent rather than by sound, and casting from labels alone gets you a young voice performing old.

What does the result look like?

The Short this page is written from is a single unbroken performance: one woman, seated in one room, telling a story straight to camera, with word-by-word captions and a quiet atmospheric bed underneath it. No cutaways, no second angle, no B-roll and no second voice anywhere in it.

It was generated in two halves because the avatar tool caps how long a single pass can be, and the halves were joined at a point where the join does not read. The viewer sees one take.

When does this not work well?

Everything rests on the performance, so there is nowhere to hide a weak voice or a drifting face. Vizard Agent tells you when the portrait or the script is fighting the format rather than adding cutaways to cover it.

How do you fix a result that came back wrong?

Say what broke the illusion for you, rather than which part you disliked. Vizard Agent keeps the cropped portrait, every voice test it ran, the generated halves and the word timings, so a new voice or a regenerated half slots straight back into the same assembly.

How does Vizard Agent compare to doing it yourself?

By hand, an avatar tool gives you one pass at a length it chooses, with whatever voice you picked from a dropdown. Vizard Agent auditions the voices by listening, works around the length cap, and assembles the parts into something that plays as one take.

By hand Vizard Agent
Voice casting Chosen from a label Generated, listened to, assessed
Length limit Accept it, or cut Split, generated in parallel, joined
Identity Drifts between passes Same optimised portrait for every part
Captions Added separately Built from the performance's own timings

Common questions

Why one continuous shot? Because cuts read as production. Vizard Agent keeps it unbroken so it reads as a person talking.

Will the face stay the same? Yes. Vizard Agent generates every part from the same cropped portrait.

How long can it be? Longer than one pass allows. Vizard Agent splits and joins rather than shortening the story.

Will the join show? It should not. Vizard Agent concatenates at a point where the picture and the voice both carry through.

How does it pick the voice? By listening. Vizard Agent generates test lines and has them assessed for age and tone.

Can I use my own voice? Yes. Vizard Agent clones it from a sample and drives the avatar with that.

Should there be music? Something subtle. Vizard Agent generates a quiet bed rather than anything with a melody.

Do I need captions? For a feed, yes. Vizard Agent builds word-by-word captions from the performance itself.

What portrait works best? Front-on, well lit, plain background. Vizard Agent crops it to the exact frame before generating.

Can it do a second episode? Yes. Vizard Agent reuses the same portrait and voice so the character stays consistent.

Why not just add B-roll? Because the whole effect is that nothing is produced. B-roll turns a person telling you something into a video about a person telling you something, which is a much weaker thing.

Does it work for a man, or a younger character? Yes. Vizard Agent runs the same casting process whatever the character is.

Does it check the result? Yes. Vizard Agent reviews frames across the finished Short for composition, identity and caption placement.