Vizard Agent

How to replace a person in a video with someone from a photo

Last updated 2026-08-16 · 6 min read

Upload the clip and a photo of the person who should appear instead, and tell Vizard Agent which person to replace. Vizard Agent picks a frame where the face is clear, tests the swap on that one frame in two versions, compares them against the original, and only then processes the clip in segments.

What is the short version?

This is a short technical job with a large consent question wrapped around it. Technically it is a test frame, a comparison and two passes over the footage. Everything difficult about it is who is in the photo and who is in the video.

  1. Go to Vizard Agent and upload the clip and the photograph.
  2. Say clearly which person in the clip is being replaced and by whom.
  3. Say that the movement, the scene and the audio should stay as they are.

What do you need before you start?

Both files and permission from both people. Vizard Agent handles the technical side by testing before committing, and the part it cannot handle is whether the person in the photograph and the person in the video have agreed to this.

What do you type into Vizard Agent?

Say who replaces whom and what must not change. Vizard Agent looks at both files itself to identify the people involved, so a short instruction naming the swap and asking for everything else to be preserved is all it needs before it starts testing on a single frame.

Prompt

Replace [the man dancing] in this video with [the man in the photo]. Keep the movement, the scene and the audio exactly as they are.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real replacement. The whole shape of the run is caution: one frame, then two versions of that frame, then a three-second test, and only then the full clip in pieces.

  1. Inspects the video and the photograph and looks at both.
  2. Extracts frames from the clip and reviews them.
  3. Checks the available tools and their prices before committing to anything.
  4. Chooses a frame where the face is clearest and previews the candidates.
  5. Uploads that reference frame and runs the swap on it in two versions.
  6. Compares both against the original and looks at the result.
  7. Runs the full clip, reads the error log when it fails, and checks the tool's duration limits.
  8. Prepares a three-second test and confirms the approach works at that length.
  9. Splits the clip into two segments, processes both, compares each against the original, and rejoins them with the original audio.

Step 5 is the habit worth copying. Two versions of a single frame cost almost nothing and told Vizard Agent whether the likeness was going to hold before any real work was spent on it.

What does the result look like?

From the run this page is written from, probed on the delivered file: 576x1024, H.264, 30fps, 3.0 seconds, AAC audio. Three seconds, vertical, with the original audio and movement intact — short, because this is a technique that holds up over seconds rather than minutes.

Vizard Agent checks the motion and the synchronisation on the finished cut rather than assuming the join between segments is clean.

It is not instant, and this category was not separately measured. Comparable work runs a median of 28 to 38 minutes end to end. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.

When does this not work well?

The technical limits are narrow and the ethical ones are narrower. Vizard Agent will do this on your own material, and putting a real person's face into footage they did not agree to appear in is not a technical decision.

How do you fix a result that came back wrong?

Say what looks wrong and where. Vizard Agent keeps the reference frame, both test versions and each processed segment separately, so trying a different source photograph or reprocessing one segment does not repeat the earlier tests, re-run the other segment, or touch the original audio at all.

How does Vizard Agent compare to doing it yourself?

By hand this is a specialist tool, a training step and a long series of trial renders — with no cost discipline at all, because every attempt is a full pass over the whole clip rather than a single frame you can look at and reject in seconds.

By hand Vizard Agent
Testing the approach Render the clip and look One frame, two versions
Hitting a tool limit Guess and retry Log read, limits checked
Long clips Process and hope Split into segments, each verified
The audio Re-attach manually Original audio kept and rejoined

Common questions

Do I need permission? Yes, from both people. This is the one requirement Vizard Agent cannot check for you.

How long can the clip be? Seconds. Vizard Agent splits longer clips into segments, and quality and cost both argue for keeping it short.

Does the audio change? No. The original sound is kept and rejoined to the processed picture.

What makes a good source photo? Front-facing, sharp, well lit, nothing across the face. It is the ceiling on the whole result.

Can I test it cheaply first? Yes, and you should. Ask Vizard Agent for a single frame or a three-second test before committing to the full clip.

Will it look real? On a short clip with a clear photo and a face-on angle, close. Turns, blur and occlusion are where it shows.

Does Vizard Agent keep the original movement? Yes. The motion, the scene and the timing are the video's; only the person changes.

Can it swap more than one person? One at a time is far more reliable. Vizard Agent tests each on a frame before processing anything.