How to put a caption on every mention of a name in a long video
Ask Vizard Agent to find every spoken instance in the word-level transcript and put the overlay on each one. The catch in an edited video is that the timings come from the source, so each mention has to be mapped onto the cut timeline — a gag that misses three out of eleven reads as a mistake.
What is the short version?
You want a caption to pop up every time someone says a particular name. Doing it by ear means scrubbing an hour of footage with your finger on the arrow key, and missing one is worse than not doing it at all, because a running gag that skips a beat looks like an editing fault rather than a joke.
- Tell Vizard Agent the word and any variant spellings of it.
- Say what should appear — a caption, a sticker, a sound.
- Say it must fire on every mention in the finished cut.
What do you need before you start?
The video and the word. Vizard Agent transcribes with word-level timings, so it does not need a list of timestamps from you, but it does need to know which variants count — a nickname, a shortened form, a name spelled two ways in the transcript.
- The video. The cut you are delivering, or the source.
- The word. And its variants.
- What appears. Text, an emoji, a graphic, a sound.
- Where. Above the speaker, at the top, in a corner.
- How long it holds. A beat, or the length of the sentence.
What do you type into Vizard Agent?
Say "every". A request that reads as an example — "put FANGIRL up when she says the name" — invites a demonstration on the first few mentions, whereas naming the requirement as exhaustive makes the completeness of the list the deliverable.
Prompt
Variants worth knowing:
- "Every single mention." The instruction that defines done.
- "Spelled two ways." Saves the misses you would never notice.
- "In the finished cut." Timings map to the edit, not the source.
What does Vizard Agent actually do?
Here is the Vizard Agent sequence on a fifteen-minute reaction edit where a recurring caption had to land on every mention of one name. The mapping step in the middle is the one that decides whether the whole thing works, and it is the step that has no equivalent when you place titles by hand.
- Transcribes both source videos with word timings.
- Finds every spoken instance of the name, in each of its spellings.
- Maps each one onto the edited timeline rather than the source timecode.
- Adds the timed overlays at those positions.
- Reads the failed filter command when the caption expression will not parse.
- Corrects the quoting in the overlay expression and renders again.
- Fixes the filter string a second time where the syntax was still wrong.
- Verifies the duration and inspects every caption moment individually.
- Uploads the captioned export and reviews the overlays on the uploaded copy.
- Quality-checks the caption timing through the opening and the outro.
Step three is the whole job. Word timings come from the transcript of the source, and the finished video has had sections cut out of it, so a mention at 9:14 in the recording is somewhere else entirely in the export. Vizard Agent translates each position through the edit rather than trusting the transcript's numbers.
Steps five to seven are a genuine fight with the tooling, kept in the list because it is the honest shape of this work. Text overlays are built as filter expressions, punctuation in a caption breaks the quoting, and Vizard Agent reads each error and corrects it rather than dropping the awkward instances.
Step eight is the check that matters to you. Not "do the captions look right" but "is there one at each of the eleven mentions", which Vizard Agent confirms mention by mention.
What does the result look like?
The same edit with a caption that fires on cue every time, holding long enough to read and gone before it becomes wallpaper. Vizard Agent tells you how many mentions it found and where they land, so you can check the count against your own memory of the footage.
When does this not work well?
Some words cannot be found reliably. A name that the transcript hears differently each time, one that overlaps with a common word, or speech under loud music will produce misses, and Vizard Agent will tell you which instances it is unsure about rather than quietly skipping them.
- Names transcribed inconsistently. Give Vizard Agent the variants.
- Words that sound like common ones. More false positives than misses.
- Heavy background music. The transcript degrades.
- Overlapping speakers. Timings blur across the crosstalk.
- Very fast repetitions. Two overlays would collide on screen.
How do you fix a result that came back wrong?
Tell Vizard Agent which mention was missed or which one fired wrongly, with a rough idea of when it happens in the cut. That is enough for it to widen the match or exclude a false positive without re-rendering the overlays that already land where you want them.
- "It missed one near the end." The transcript is re-checked for that spelling.
- "It fired on the wrong word." That match is excluded by context.
- "They are on screen too long." The hold is shortened across the set.
- "One covers her face." The position is moved for all of them, or just that one.
How does Vizard Agent compare to doing it yourself?
By hand you scrub, you place a title, you scrub again. Fifteen minutes of footage with eleven mentions is an hour of work, and the eleventh is the one you miss — which is the one a commenter will point out under the video.
| By hand | Vizard Agent | |
|---|---|---|
| Finding mentions | Scrubbing and listening | Word-level transcript search |
| Variants | Whatever you remember | Every spelling you name |
| After an edit | Timings drift | Mapped onto the cut timeline |
| Completeness | Hoped for | Checked mention by mention |
| Consistency | Drifts across the video | One treatment, applied to all |
Common questions
Can it do a sound effect instead of a caption? Yes. Vizard Agent can place a sound on each mention as easily as a graphic.
Can it count them for me? Yes. Vizard Agent reports how many mentions it found and where.
What about two different words? Give both, and Vizard Agent can give them different treatments.
Will it work on an hour-long video? Yes, though transcription and rendering take proportionally longer.
Can the caption be a picture? Yes, an image or a sticker can be timed the same way.
What if the name is misheard? Tell Vizard Agent the variants and it matches those too.
Does it work in other languages? Yes. The transcript carries word timings in the spoken language.
Can it skip the first one? Yes, if you want the gag to build. Say which mentions to include.
What if I already have a subtitle file? Give it to Vizard Agent and it works from your timings rather than its own.
Can the overlay move out of the way of faces? Yes. Say so, and the position is checked against the frame.
Does the original audio change? No, unless you asked for a sound effect on each mention.
Will the captions survive a re-edit? Tell Vizard Agent about the re-edit and the positions are mapped again.
How long should each one hold? About a second reads well. Vizard Agent will match whatever you specify.
Can it do this on several episodes? Yes. Once the treatment is agreed, Vizard Agent repeats it across the series.