How to make two photos talk to each other in one video
Upload both photographs and write what each person says. Vizard Agent casts a voice for each, generates the two speaking shots separately, checks the frames and the audio of each before joining them, and builds the captions from a transcript of the assembled cut rather than the script.
What is the short version?
A two-photo dialogue is two separate generations joined by a transition, not one video with two people in it. Each shot has to hold on its own, and the handover between them is what makes it read as a conversation rather than two clips.
- Go to Vizard Agent and upload both photographs.
- Write what each person says, in order.
- Describe each voice and how the shots should hand over.
What do you need before you start?
Two clear photographs and both halves of the dialogue. Vizard Agent casts the voices, generates each shot and handles the join, so the material requirement is small. Writing the handover explicitly is worth doing, because how the second person enters is what makes the two shots feel connected.
- Two photographs. One per character, front-facing and clear.
- Both parts of the dialogue. Written out exactly.
- The voices. Age, warmth, accent, language.
- The handover. How the second character comes into frame.
- Your budget. Say it; Vizard Agent will plan the shots to fit.
What do you type into Vizard Agent?
Write it as a two-shot script. Vizard Agent generates one shot per speaker, so describing what happens in each — who is on camera, what they say, how the frame moves to the other — maps directly onto how the video gets built.
Prompt
Variants worth knowing:
- A selfie handover. The first character turns the camera towards the second.
- A cloned voice. If you have a recording of the person.
- A stated credit budget. Vizard Agent plans the shot count to fit it.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real two-photo dialogue. Notice that each shot is checked twice — once for the picture and once for what the audio actually says — before the two are joined.
- Opens both uploaded photographs and checks the supplied audio and available tools.
- Searches for a voice matching the description across two passes.
- Clones the supplied voice for the first character's line and generates both characters' audio.
- Checks the generated audio and its timestamps and analyses the result.
- Generates the first shot and probes its specifications, extracting a frame sequence.
- Reviews the first shot's frames and checks what its audio actually says.
- Generates the second shot and reviews both its frames and its speech the same way.
- Sources warm background music and joins the two shots losslessly into a rough cut.
- Transcribes the assembled cut for word-level timings, generates animated captions, checks the installed typefaces for the language and renders test frames until the subtitles display correctly.
Step 9's opening decision is the one to copy. Building captions from the original script assumes the generated speech matches it exactly, and transcribing the assembled cut instead means the timings come from what was actually said.
What does the result look like?
From the run this page is written from, probed on the delivered file: 720x1280, H.264, 24fps, 15.32 seconds, AAC audio. Vertical, fifteen seconds, two photographs brought to life in sequence with a handover between them, both voices cast to description, animated captions and music underneath.
Fifteen seconds is the right size for an exchange of two lines. This format works because it is brief and direct; a longer dialogue starts to expose that neither shot is really moving.
When does this not work well?
Making two real people appear to speak to each other on camera is by some distance the most sensitive thing described anywhere in this section. Vizard Agent will build it carefully and well, and every question about whether it should be built at all sits with you.
- Use only photographs you have the right to use. A likeness belongs to the person in it.
- Never script a real person saying something they did not. That is the line this must not cross.
- For someone who has died, the family decides. A memorial exchange is theirs to want.
- Do not present it as real footage. These are animated photographs, and saying so matters.
- Two shots is the practical limit. A longer conversation needs a different approach entirely.
How do you fix a result that came back wrong?
Name the shot or the line. Vizard Agent keeps both generated shots, the voice takes, the assembled transcript and the caption tests, so re-recording one line or regenerating one shot is a targeted rebuild rather than starting over.
- "The second voice is too young." Another candidate, generated and checked.
- "The handover is abrupt." Regenerated with a longer move into frame.
- "The subtitles do not match the audio." Rebuilt from the assembled cut's own transcript.
How does Vizard Agent compare to doing it yourself?
By hand this means running each photograph through a lip-sync tool separately, then cutting them together and hoping the join reads as a conversation. The captions are the part that goes wrong quietly, because typing them from your script produces text that drifts against speech the generator paced differently.
| By hand | Vizard Agent | |
|---|---|---|
| Each shot | Generate and accept | Frames and audio both checked per shot |
| The handover | A hard cut | Generated as a described move into frame |
| Captions | Typed from the script | Transcribed from the assembled cut |
| Typeface for the language | Discover at render | Tested on frames until it displays |
Common questions
What is this format actually for? Most often a family message: someone recording words for a relative, or a keepsake built around a photograph. It works because it is short, personal and clearly presented as animated rather than filmed.
Can Vizard Agent use my own voice recording? Yes. Vizard Agent clones from the audio you supply rather than generating a voice from scratch, which matters when the video is for someone who knows what that person sounds like.
How many characters can talk? Two is the practical limit here, since each speaker is generated as their own shot. A longer conversation needs a different approach entirely.
Will the faces distort? Vizard Agent extracts frames from each generated shot and checks before joining them.
Can Vizard Agent work to a budget? Yes. State it and Vizard Agent plans the shot count and the generation choices to fit.
Why transcribe the cut instead of using the script? Because generated speech does not follow a script's punctuation exactly, and captions built from the written lines drift against it. Transcribing what was actually said gives timings that match the audio to the word.
Will it work in my language? Yes, and Vizard Agent tests the typeface renders it before finalising the captions.
Does Vizard Agent add music? Yes. Vizard Agent sources a bed that matches the tone of the exchange rather than defaulting to one.
Does Vizard Agent check the rendered captions? Yes. It rendered test frames repeatedly here until the subtitles displayed correctly.