Vizard Agent

How to make two photos talk to each other in one video

Last updated 2026-08-24 · 7 min read

Upload both photographs and write what each person says. Vizard Agent casts a voice for each, generates the two speaking shots separately, checks the frames and the audio of each before joining them, and builds the captions from a transcript of the assembled cut rather than the script.

What is the short version?

A two-photo dialogue is two separate generations joined by a transition, not one video with two people in it. Each shot has to hold on its own, and the handover between them is what makes it read as a conversation rather than two clips.

  1. Go to Vizard Agent and upload both photographs.
  2. Write what each person says, in order.
  3. Describe each voice and how the shots should hand over.

What do you need before you start?

Two clear photographs and both halves of the dialogue. Vizard Agent casts the voices, generates each shot and handles the join, so the material requirement is small. Writing the handover explicitly is worth doing, because how the second person enters is what makes the two shots feel connected.

What do you type into Vizard Agent?

Write it as a two-shot script. Vizard Agent generates one shot per speaker, so describing what happens in each — who is on camera, what they say, how the frame moves to the other — maps directly onto how the video gets built.

Prompt

Shot one: the person in photo one says "[line]" to camera. Then walk into frame and bring in the person from photo two, who says "[line]". Voices: [descriptions]. Around [15] seconds.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real two-photo dialogue. Notice that each shot is checked twice — once for the picture and once for what the audio actually says — before the two are joined.

  1. Opens both uploaded photographs and checks the supplied audio and available tools.
  2. Searches for a voice matching the description across two passes.
  3. Clones the supplied voice for the first character's line and generates both characters' audio.
  4. Checks the generated audio and its timestamps and analyses the result.
  5. Generates the first shot and probes its specifications, extracting a frame sequence.
  6. Reviews the first shot's frames and checks what its audio actually says.
  7. Generates the second shot and reviews both its frames and its speech the same way.
  8. Sources warm background music and joins the two shots losslessly into a rough cut.
  9. Transcribes the assembled cut for word-level timings, generates animated captions, checks the installed typefaces for the language and renders test frames until the subtitles display correctly.

Step 9's opening decision is the one to copy. Building captions from the original script assumes the generated speech matches it exactly, and transcribing the assembled cut instead means the timings come from what was actually said.

What does the result look like?

From the run this page is written from, probed on the delivered file: 720x1280, H.264, 24fps, 15.32 seconds, AAC audio. Vertical, fifteen seconds, two photographs brought to life in sequence with a handover between them, both voices cast to description, animated captions and music underneath.

Fifteen seconds is the right size for an exchange of two lines. This format works because it is brief and direct; a longer dialogue starts to expose that neither shot is really moving.

When does this not work well?

Making two real people appear to speak to each other on camera is by some distance the most sensitive thing described anywhere in this section. Vizard Agent will build it carefully and well, and every question about whether it should be built at all sits with you.

How do you fix a result that came back wrong?

Name the shot or the line. Vizard Agent keeps both generated shots, the voice takes, the assembled transcript and the caption tests, so re-recording one line or regenerating one shot is a targeted rebuild rather than starting over.

How does Vizard Agent compare to doing it yourself?

By hand this means running each photograph through a lip-sync tool separately, then cutting them together and hoping the join reads as a conversation. The captions are the part that goes wrong quietly, because typing them from your script produces text that drifts against speech the generator paced differently.

By hand Vizard Agent
Each shot Generate and accept Frames and audio both checked per shot
The handover A hard cut Generated as a described move into frame
Captions Typed from the script Transcribed from the assembled cut
Typeface for the language Discover at render Tested on frames until it displays

Common questions

What is this format actually for? Most often a family message: someone recording words for a relative, or a keepsake built around a photograph. It works because it is short, personal and clearly presented as animated rather than filmed.

Can Vizard Agent use my own voice recording? Yes. Vizard Agent clones from the audio you supply rather than generating a voice from scratch, which matters when the video is for someone who knows what that person sounds like.

How many characters can talk? Two is the practical limit here, since each speaker is generated as their own shot. A longer conversation needs a different approach entirely.

Will the faces distort? Vizard Agent extracts frames from each generated shot and checks before joining them.

Can Vizard Agent work to a budget? Yes. State it and Vizard Agent plans the shot count and the generation choices to fit.

Why transcribe the cut instead of using the script? Because generated speech does not follow a script's punctuation exactly, and captions built from the written lines drift against it. Transcribing what was actually said gives timings that match the audio to the word.

Will it work in my language? Yes, and Vizard Agent tests the typeface renders it before finalising the captions.

Does Vizard Agent add music? Yes. Vizard Agent sources a bed that matches the tone of the exchange rather than defaulting to one.

Does Vizard Agent check the rendered captions? Yes. It rendered test frames repeatedly here until the subtitles displayed correctly.