Vizard Agent

How to animate one person in a photo and leave everyone else still

Last updated 2026-09-05 · 8 min read

Tell Vizard Agent which person in the photograph should move and that nothing else may change. It pins the first and last frame of each generated segment back to your original image, measures those frames for drift, and reviews any replacement shot before merging it so a section with good movement is never swapped for one with worse.

What is the short version?

Animating a photograph tends to animate everything in it: faces shift, clothes change, the camera decides to move. Keeping one person alive and everybody else exactly as photographed is a matter of anchoring each generated second back to the original picture.

  1. Go to Vizard Agent with the photograph.
  2. Say who should move and that the rest must not change.
  3. Ask for the frames to be checked against the original.

What do you need before you start?

The photograph and a clear list of what must stay. That list is doing the real work — the faces, the clothing, the setting and the camera angle are the things a generator will quietly improve unless you say they are fixed.

What do you type into Vizard Agent?

Write the constraints as a list of things that stay the same. It reads as over-specified and it is not: each item you leave out is an item that may drift, and drift in a photograph of real people is the thing everyone notices.

Prompt

Turn this photo into a realistic, smooth video while keeping the original picture, the scene composition, the people, their clothes and the place exactly as they are. Do not change anyone's appearance, do not add people, and do not move the camera much. Only the man in the middle should move.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real photograph that had to become a thirty-second video with only one of its subjects moving. The short test in the middle of the list is the thing that makes everything after it hold together.

  1. Examines the photograph and its scene composition first.
  2. Checks the options for animating a still, and their limits.
  3. Picks the approach that holds the scene steadiest.
  4. Turns the reference image into a first clip with quiet movement.
  5. Extracts frames to review how steady the scene stayed.
  6. Tests a shorter clip with the start and end pinned to the original image.
  7. Measures the first and last frames for any large change.
  8. Generates several short clips in parallel, all anchored to the photo.
  9. Joins them and lays the original audio across the whole thing.
  10. Reviews the joins between clips in a single grid.
  11. Generates a replacement section with clearer movement for the subject.
  12. Reviews that replacement before merging it, so nothing good is lost.
  13. Swaps in the better section and re-mixes the audio in full.
  14. Smooths the jumps between the short clips at a fixed duration.

Step six is the technique. A generated clip drifts furthest at its ends, so pinning both ends to the original photograph keeps every segment returning to the picture you actually supplied, and the drift never accumulates across the whole video.

Step twelve is the habit that saves work. A regenerated section is not automatically better than the one it replaces, so it gets looked at first — otherwise you spend a round improving the movement and discover the faces got worse.

What does the result look like?

The photograph, alive in exactly one place: your chosen subject moving naturally, everyone around them sitting exactly as they were photographed, the same clothes, the same room, the same framing Vizard Agent started from, with your audio running the full length.

Held next to the original still, only the intended thing has changed.

When does this not work well?

Anchoring costs movement. The more strictly each segment is pinned to the photograph, the less anyone in it can do, so a large gesture or a walk across the frame is difficult to get without letting other things drift as well.

How do you fix a result that came back wrong?

Say who moved that should not have, or who did not move that should. Vizard Agent keeps every segment and its anchor frames, so a fix replaces one section rather than regenerating the whole video from the photograph again.

How does Vizard Agent compare to doing it yourself?

The usual approach is one long generation from the photograph, which drifts steadily from beginning to end — by the last seconds the faces are approximate and the room has quietly redecorated. Short anchored segments cost more calls and hold the picture.

By hand Vizard Agent
Structure One long generation Short segments, each anchored
Drift Accumulates to the end Reset at every segment
Checking it Watched once First and last frames measured
Replacing a section Accepted on faith Reviewed before merging
The audio Re-timed after each attempt Laid across a fixed duration

Common questions

Can only one person move? Yes. Name them and Vizard Agent keeps the others still.

Will faces stay the same? That is what the anchoring is for. Vizard Agent measures it.

Can I set an exact length? Yes. Vizard Agent holds the duration and fits the audio to it.

Will the camera move? Only if you ask. Vizard Agent keeps it still by default.

Can it add ambient sound? Yes. Tell Vizard Agent what should be audible and how distant.

What if a section looks worse? Vizard Agent reviews replacements before merging them.

Can it animate an old photograph? Yes, though Vizard Agent finds damage and grain make anchoring harder.

Will the joins be visible? Vizard Agent smooths them and checks each boundary.

Can it make the person speak? Yes. Vizard Agent does that as a separate step with a voice track.

How long can it run? As long as you like, though Vizard Agent notes drift risk grows with length.

Can I supply my own audio? Yes. Vizard Agent lays it across the whole video.

What about a group photo? Harder, because Vizard Agent has to hold every face at once.

Can it move the subject a lot? Some. For Vizard Agent, large movement and strict anchoring pull against each other.

Why anchor each segment rather than generate once? Because a single long generation has nothing to return to, so every second is built on the last one's mistakes.