Vizard Agent

How to make a still photo speak a recorded message

Last updated 2026-08-24 · 6 min read

Upload the photo and give Vizard Agent the message. Vizard Agent searches for a voice with the right age and warmth, records the words, drives the face from that recording so the mouth matches, checks the result frame by frame for distortion, and cuts it together with music and captions.

What is the short version?

A photograph can be made to deliver a message in a voice you choose. The technical part is ordinary now; what decides whether the result is worth keeping is the voice and the face holding steady, and both of those are checked rather than assumed.

  1. Go to Vizard Agent and upload the photo.
  2. Give it the exact words to be spoken.
  3. Describe the voice — the age, the warmth, the language.

What do you need before you start?

A photo and the words. Vizard Agent finds the voice and drives the face itself, so nothing needs preparing beyond those two, and describing the voice carefully is what separates a result that moves someone from one that feels mechanical.

What do you type into Vizard Agent?

Describe the voice in human terms. Vizard Agent searches by that description rather than by technical settings, so "an elderly woman, gentle and kind" narrows the search far more usefully than any parameter would, and it will play candidates back before committing.

Prompt

Make the person in this photo speak this message: "[the words]". The voice should be [an elderly woman, gentle and warm], in [language]. Keep the face natural, no distortion.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real message video. The long middle section is not the talking head at all — it is removing a small watermark that sat across the subject's chest, which took a dozen attempts to do without damaging the picture.

  1. Opens both photos and analyses the supplied audio.
  2. Searches for an elderly female voice with the right warmth in the target language.
  3. Generates the narration and evaluates its quality and tone before using it.
  4. Drives the first photo from the audio and extracts frames to check the face.
  5. Generates the second speaking shot and reviews its frames the same way.
  6. Generates a music bed and transition sound effects, then measures every track in LUFS.
  7. Tests the transition between the two people and inspects the crossover frames.
  8. Locates a stray watermark by pixel coordinates, tries feathered repair, then precise repair, then repair from the original image, checking the frame after every attempt.
  9. Extracts word-level timings, builds animated captions in a font that renders the language, renders the full cut, then refines the camera move between the two shots and re-renders.

Step 8 is the part nobody plans for and everybody hits. A repair that is slightly too aggressive smears the face next to it, and the only way to land it is to look at the frame after each attempt rather than trusting the filter.

What does the result look like?

From the run this page is written from, probed on the delivered file: 720x1280, H.264, 24fps, 13.67 seconds, AAC audio. Vertical, under fourteen seconds, two photographs brought to life in sequence, the message spoken in a chosen voice, with music and animated captions in the message's own language.

Fourteen seconds is the right length for this. A spoken message from a photograph works because it is short and direct; stretched past half a minute the effect starts to draw attention to itself rather than to what is being said.

When does this not work well?

This is the most sensitive kind of video described anywhere on this site, and almost everything about whether it should be made at all sits outside the edit itself. Vizard Agent will produce it carefully and well; the judgement about whether to make it is entirely yours.

How do you fix a result that came back wrong?

Say what feels off. Vizard Agent keeps the voice candidates, the generated audio, the word timings, every repair attempt and the rendered shots, so changing the voice or the wording is a re-record rather than a rebuild from the photograph.

How does Vizard Agent compare to doing it yourself?

By hand this is picking a synthetic voice from a dropdown menu, running the photo through a lip-sync tool once, and then accepting whatever comes out of it, including the watermark still sitting across the subject's chest. Vizard Agent checks each stage instead of trusting any of them.

By hand Vizard Agent
The voice Pick from a list Searched by description, then auditioned
The face Trust the tool Frames extracted and checked for distortion
A blemish in the photo Live with it Located by pixel, repaired, verified each pass
Captions Type and nudge Built from word-level timings

Common questions

What kind of photo works best? A clear, front-facing face at reasonable resolution. Angled or blurred photos distort.

Can I use my own voice recording? Yes. Upload the audio and Vizard Agent drives the face from it rather than generating a voice.

Can two people appear? Yes. Vizard Agent brought a second photo into frame on this run and handed the message across.

Will it look distorted? Vizard Agent extracts frames and checks the face specifically for stretching before it accepts the shot.

Can it work in my language? Yes, for the voice and the captions, and it checks the font can render the script.

How long should the message be? Ten to twenty seconds. Short and direct is what makes this format land.

Is it obvious that it is generated? To most viewers, yes, and that is fine. What matters is not presenting it as real footage of the person speaking.

Does Vizard Agent check the mix? Yes. Vizard Agent measures the voice, the music and the effects in LUFS and inspects the waveform of the finished file before delivery.

Can it remove something from the photo? Yes. Vizard Agent located and repaired a watermark on this run, checking the frame after each attempt.