Vizard Agent

How to land a photo on the exact line that mentions the person

Last updated 2026-09-22 · 8 min read

Tell Vizard Agent which photo belongs to which mention. A card that appears while she is talking about her father is doing something completely different from one that appears when she says the word — and the second is what makes a personal reel feel edited rather than illustrated.

What is the short version?

You have a finished reel with subtitles on it, and two photographs that belong to two specific sentences inside it. Everything else about the video is already right, so what you want from Vizard Agent is an addition rather than an edit.

  1. Tell Vizard Agent which photo goes with which mention.
  2. Say the existing subtitles stay exactly as they are.
  3. Ask it to check the cards do not collide with the captions.

What do you need before you start?

The reel, the photos and the two moments. Vizard Agent reads the subtitles and the speech to find the lines, so you can describe the moments in words — "when she mentions her father" — rather than hunting for timestamps.

What do you type into Vizard Agent?

Describe the mentions rather than the times. The transcript is what Vizard Agent matches against, so "when she mentions her father" is not vague to it — it is a precise instruction that survives any change to the edit.

Prompt

Take this reel and keep the same subtitles. When she mentions her father, put up the photo of her and her dad. When she mentions her mother, put up the family photo. Keep everything else as it is, and add anything small that makes it land better.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on a personal reel that needed two photo cards adding to it. Most of the care in the list goes into where on the frame those cards sit, rather than into working out the moment at which they should appear.

  1. Inspects the reel and the photos.
  2. Views the reel's frames to see the existing treatment.
  3. Looks at the photos alongside a frame of the reel to judge the fit.
  4. Checks exactly where the subtitle pill sits in the frame.
  5. Views the subtitle seam close up, because a card placed over it would look like a mistake.
  6. Measures the layout and the dialogue loudness.
  7. Generates music and measures it against the dialogue level.
  8. Builds the photo card frames and checks the music underneath them.
  9. Renders the reel with the cards and the music.
  10. Checks the rendered frames and the loudness.
  11. Reviews each card moment, then inspects the card detail at full size.
  12. Fetches the photographs from their original posts when better copies are needed.

Step four is the step that separates this from dropping a picture on the timeline. An existing reel already has a subtitle position, and a photo card that covers it either hides the words or fights them for the same strip of screen.

Step eleven matters because a card is only on screen for a second. Checking it at full size is how you find that the photo is cropped through somebody's head, which is invisible at playback speed and obvious to the family.

Step twelve is the practical footnote. The best copy of a photo is often the original post rather than the screenshot you were sent, and it is worth going back for it.

What does the result look like?

The same reel from Vizard Agent with two photographs appearing exactly on the lines that mention those people, sitting clear of the subtitles, held long enough to actually look at, and with music underneath that never competes with her voice.

When does this not work well?

Some mentions cannot carry a card. If the line is half a second long, if two people are named in one sentence, or if the photo is low resolution, the card is either a flash or a distraction.

How do you fix a result that came back wrong?

Say which card and what is wrong with it — its timing, its size or its position. Vizard Agent treats each card as its own element, so one can be moved without disturbing the other or the subtitles underneath.

How does Vizard Agent compare to doing it yourself?

By hand this is easy and fiddly: scrub to the word, drop the photo, nudge it off the captions, repeat. It goes wrong in the small ways — a card half a second late, a crop through a face — which are exactly the things nobody notices until the person in the photo watches it.

By hand Vizard Agent
Timing Scrubbed to by ear Matched to the word in the transcript
Position Nudged off the captions Measured against the caption strip
The crop Checked at preview size Checked at full size
Subtitles Often rebuilt Kept exactly as they are
Music Added and levelled by ear Measured against the dialogue

Common questions

Can it find the mention itself? Yes. Vizard Agent matches your description against the transcript.

Will my subtitles change? No, if you say to keep them — Vizard Agent works around them.

How long should a photo stay up? Vizard Agent holds it for the length of the line that names it.

Can it be a full-screen card? Yes, or a corner inset — Vizard Agent will suggest which suits the layout.

What if I have several photos of the same person? Vizard Agent cycles them across the mentions, or picks the clearest one.

Can it animate the card? Yes. Vizard Agent uses a gentle push rather than a static drop.

Does the reel get longer? Only if you want it to. Cards can sit over the existing picture instead.

Can it add music at the same time? Yes, and Vizard Agent measures it against her voice so it stays underneath.

What about photos from a social post? Vizard Agent can fetch the original if you give it the link.

Will it work with burnt-in captions? Yes. Vizard Agent measures where they sit and keeps the card clear.

Can it caption the photo? Yes — a name and a year is a common addition.

What if the mention is in a different language? The transcript still finds it, whatever language it is in.

Can it do this across a series? Yes, with the same card treatment each time.

How do I check it? Watch it once for the words and once for the pictures; they should agree.