Vizard Agent

How to present a video as a 2D avatar instead of on camera

Last updated 2026-09-22 · 8 min read

Send Vizard Agent your sprite sheet, a sample of your voice and the footage you want covered. It clones the voice for the narration, locates each pose on the sheet by its silhouette, re-crops them cleanly, and switches expressions in time with what is being said.

What is the short version?

You present as a character rather than on camera, and you want a proper edit built around that character — covering footage, captions, and narration in your own voice. Vizard Agent treats the sprites as the host layer of the video rather than as decoration laid on top of it.

  1. Send Vizard Agent the sprites, a voice sample and the clips.
  2. Let it re-crop the poses from the sheet.
  3. Check the captions do not land on text already in your footage.

What do you need before you start?

The sprite sheet or the individual poses, a clean recording of your voice, and the clips you want to talk over. A minute of clear speech is enough for the voice; the sprites matter more, because a sheet with poses that touch each other is harder to separate than one with space between them.

Say what each pose is for if it is not obvious. Neutral, talking, surprised, annoyed — naming them means Vizard Agent can switch on meaning rather than guessing from the drawing.

What do you type into Vizard Agent?

Say what you are and hand over the three things. The word for this format carries a lot of implicit specification with it — where the host sits, that the poses switch rather than animate — so using it saves you explaining the layout from scratch.

I am a PNGtuber. Here is my voice to clone, my sprites, and the clips I want used. Make a video about this topic.

That is close to the real brief. What it leaves unsaid — where the character sits, how big, when the expressions change — Vizard Agent proposes and you correct, which is faster than specifying it blind.

What does Vizard Agent actually do?

It gets the sprites right before it builds anything, because a badly cropped host is visible in every frame. Poses on a sheet are located by their visible silhouette rather than by an even grid, then re-cropped tight so no neighbouring pose bleeds into the frame.

In that session the work ran:

Then two problems that are specific to this format: the narration was split into two sections to stay inside the voice engine's length limit, and the captions had to be moved because the source clips already had text burnt into them.

What does the result look like?

Your character hosting, in your voice, over your footage, with the expression changing at the points where it should. The sprites sit cleanly with no halo and no fragment of the adjacent pose, and the captions are in a band of their own.

The script is checked as well as the edit. Because the narration here is written rather than transcribed from something you said, the facts in it are researched against real sources rather than generated, which matters more in this format than on a talking-head cut where you said the words yourself and stand behind them.

When does this not work well?

When the sprite sheet is tightly packed. Poses drawn touching each other cannot be separated cleanly by silhouette, and the crops will either clip the character or include a neighbour. Individual files solve it completely if you have them.

It also does not do real lip sync. The format is pose switching, not mouth shapes, and a brief that expects visible mouth movement per phoneme wants a different technique entirely.

And a voice sample with background noise clones the noise along with the voice. That is worth re-recording before anything else happens, because every second of the finished narration carries it, and there is no pass afterwards that removes it cleanly.

How do you fix a result that came back wrong?

If a pose looks wrong, name that pose and say how it is wrong — clipped, haloed, or showing a piece of the one next to it. Vizard Agent re-crops that single pose from the sheet rather than redoing the whole extraction and disturbing the others.

If the expressions switch at the wrong moments, give a line rather than a timestamp. The switches are anchored to the transcript, so "surprised should be on the line about the data" moves it precisely.

If the captions collide with text in your footage, say which clip. Vizard Agent moves the caption band for that section rather than for the whole video, so it stays consistent everywhere else.

How does Vizard Agent compare to doing it yourself?

The usual setup is a live tool that swaps sprites on your microphone input while you record. That is fine for streaming and it produces a take you cannot edit — the expression changes are baked in at the moment you spoke.

Building it from a script means the switches are editable and the narration can be re-recorded without redoing the host. The trade is that it is not live. For a written, researched piece rather than a stream, that is usually the right way round.

Common questions

Do I need separate pose files? No, a sheet works. Separate files crop more reliably, though.

How much voice do you need? About a minute of clean speech. Vizard Agent will say if the sample is too noisy.

Can it do mouth movement? Not per phoneme. This format switches poses rather than animating lips.

How many expressions can I use? As many as you supply. Vizard Agent switches on meaning from the transcript.

Can the character move? Subtly, yes. Vizard Agent can add a small bob or scale on emphasis.

Where does the host sit? Wherever you say. Vizard Agent proposes a corner if you do not.

Can it write the script? Yes, researched against real sources rather than invented.

What if my narration is long? Vizard Agent splits it into sections within the voice engine's limit and joins them.

Will the joins be audible? They should not be. Vizard Agent checks the seams on the rendered audio.

Can I use a transparent background? Yes, that is how the sprites are composited over your footage.

What about captions? Included, and Vizard Agent keeps them clear of text already in your clips.

Can it match my channel's look? Yes. Send an earlier video and Vizard Agent matches the treatment.

Is landscape or vertical better? Whichever your channel uses. Say which and Vizard Agent frames for it.

Can I swap the sprites later? Yes. The host is its own layer, so Vizard Agent replaces it without recutting the video.

Can two characters appear at once? Yes, if you supply both sets. Vizard Agent switches each on its own lines.

Does the background music need to duck? Yes, under the narration. Vizard Agent measures it rather than setting it by ear.