How to animate one person in a photo and leave everyone else still
Tell Vizard Agent which person in the photograph should move and that nothing else may change. It pins the first and last frame of each generated segment back to your original image, measures those frames for drift, and reviews any replacement shot before merging it so a section with good movement is never swapped for one with worse.
What is the short version?
Animating a photograph tends to animate everything in it: faces shift, clothes change, the camera decides to move. Keeping one person alive and everybody else exactly as photographed is a matter of anchoring each generated second back to the original picture.
- Go to Vizard Agent with the photograph.
- Say who should move and that the rest must not change.
- Ask for the frames to be checked against the original.
What do you need before you start?
The photograph and a clear list of what must stay. That list is doing the real work — the faces, the clothing, the setting and the camera angle are the things a generator will quietly improve unless you say they are fixed.
- The photo. At the best resolution you have.
- Who moves. By position if you cannot name them.
- What must not change. Faces, clothes, place, framing.
- How long it should run. And whether that is exact.
- Any audio that has to sit under the whole thing.
What do you type into Vizard Agent?
Write the constraints as a list of things that stay the same. It reads as over-specified and it is not: each item you leave out is an item that may drift, and drift in a photograph of real people is the thing everyone notices.
Prompt
Variants worth knowing:
- "Do not add people." Generators like a fuller frame.
- "Do not move the camera." Otherwise it will.
- "Only this person moves." Names the exception.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real photograph that had to become a thirty-second video with only one of its subjects moving. The short test in the middle of the list is the thing that makes everything after it hold together.
- Examines the photograph and its scene composition first.
- Checks the options for animating a still, and their limits.
- Picks the approach that holds the scene steadiest.
- Turns the reference image into a first clip with quiet movement.
- Extracts frames to review how steady the scene stayed.
- Tests a shorter clip with the start and end pinned to the original image.
- Measures the first and last frames for any large change.
- Generates several short clips in parallel, all anchored to the photo.
- Joins them and lays the original audio across the whole thing.
- Reviews the joins between clips in a single grid.
- Generates a replacement section with clearer movement for the subject.
- Reviews that replacement before merging it, so nothing good is lost.
- Swaps in the better section and re-mixes the audio in full.
- Smooths the jumps between the short clips at a fixed duration.
Step six is the technique. A generated clip drifts furthest at its ends, so pinning both ends to the original photograph keeps every segment returning to the picture you actually supplied, and the drift never accumulates across the whole video.
Step twelve is the habit that saves work. A regenerated section is not automatically better than the one it replaces, so it gets looked at first — otherwise you spend a round improving the movement and discover the faces got worse.
What does the result look like?
The photograph, alive in exactly one place: your chosen subject moving naturally, everyone around them sitting exactly as they were photographed, the same clothes, the same room, the same framing Vizard Agent started from, with your audio running the full length.
Held next to the original still, only the intended thing has changed.
When does this not work well?
Anchoring costs movement. The more strictly each segment is pinned to the photograph, the less anyone in it can do, so a large gesture or a walk across the frame is difficult to get without letting other things drift as well.
- Big movements. They fight the anchoring.
- Long runtimes. More segments, more chances to drift.
- Crowded photographs. More faces to keep still.
- Low resolution. Detail invents itself when enlarged.
- A moving camera. That changes everything at once.
How do you fix a result that came back wrong?
Say who moved that should not have, or who did not move that should. Vizard Agent keeps every segment and its anchor frames, so a fix replaces one section rather than regenerating the whole video from the photograph again.
- "The middle one is still frozen." That section regenerated.
- "Someone's face changed." Re-anchored to the original.
- "There is a jump at a join." Smoothed at that boundary.
- "The audio drifted." Re-laid across the fixed duration.
How does Vizard Agent compare to doing it yourself?
The usual approach is one long generation from the photograph, which drifts steadily from beginning to end — by the last seconds the faces are approximate and the room has quietly redecorated. Short anchored segments cost more calls and hold the picture.
| By hand | Vizard Agent | |
|---|---|---|
| Structure | One long generation | Short segments, each anchored |
| Drift | Accumulates to the end | Reset at every segment |
| Checking it | Watched once | First and last frames measured |
| Replacing a section | Accepted on faith | Reviewed before merging |
| The audio | Re-timed after each attempt | Laid across a fixed duration |
Common questions
Can only one person move? Yes. Name them and Vizard Agent keeps the others still.
Will faces stay the same? That is what the anchoring is for. Vizard Agent measures it.
Can I set an exact length? Yes. Vizard Agent holds the duration and fits the audio to it.
Will the camera move? Only if you ask. Vizard Agent keeps it still by default.
Can it add ambient sound? Yes. Tell Vizard Agent what should be audible and how distant.
What if a section looks worse? Vizard Agent reviews replacements before merging them.
Can it animate an old photograph? Yes, though Vizard Agent finds damage and grain make anchoring harder.
Will the joins be visible? Vizard Agent smooths them and checks each boundary.
Can it make the person speak? Yes. Vizard Agent does that as a separate step with a voice track.
How long can it run? As long as you like, though Vizard Agent notes drift risk grows with length.
Can I supply my own audio? Yes. Vizard Agent lays it across the whole video.
What about a group photo? Harder, because Vizard Agent has to hold every face at once.
Can it move the subject a lot? Some. For Vizard Agent, large movement and strict anchoring pull against each other.
Why anchor each segment rather than generate once? Because a single long generation has nothing to return to, so every second is built on the last one's mistakes.