Vizard Agent

How to highlight only the numbers your speaker says in the captions

Last updated 2026-09-10 · 8 min read

Tell Vizard Agent which phrases to emphasise rather than asking for emphasis in general. It reads the word timings to find exactly where those figures are spoken, groups the captions into three or four words a screen, recolours only the phrases you named, and checks each highlight on a rendered frame so the colour lands on the whole amount and not half of it.

What is the short version?

Caption presets that highlight the "active" word colour everything in turn, which is a rhythm effect rather than an emphasis one. If two numbers are the point of your video, those two should be the only coloured things in it.

  1. Go to Vizard Agent with the cut.
  2. Name the exact phrases to highlight.
  3. Ask it to check each highlight on a frame.

What do you need before you start?

The cut and the phrases, written out as they are said. Vizard Agent matches against the spoken words, so "the price" is not enough — the figure as spoken is what it needs, including whether the speaker says the currency and whether they round it.

What do you type into Vizard Agent?

Name the phrases and the style in one go. Asking for "emphasis on the important bits" hands the judgement back, and what you get is a preset that animates every word rather than the two that carry the argument.

Prompt

Caption this in grouped lines of three or four words, plain white, and recolour only the two spoken dollar amounts in yellow. Check each highlight on a rendered frame afterwards so I can see the colour covers the whole amount and the captions clear her face.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the sequence on a real vertical short whose whole point was two figures. Notice the work that happens before any colour is applied: the phrases are located in the word timings first, so the highlight is attached to speech rather than to a guess.

  1. Reads the source word timings.
  2. Confirms the spoken text around each figure.
  3. Rebuilds a retimed transcript for the cut itself.
  4. Groups the captions into three or four words a screen.
  5. Generates the caption track in the plain style.
  6. Recolours only the named phrases.
  7. Renders and pulls frames at each highlight.
  8. Verifies the first amount on its own frame.
  9. Verifies the second amount on its own frame.
  10. Repositions the captions if they cover a face.
  11. Re-checks the highlights after the reposition.

Step three is the step that gets skipped and causes most of the misses. Word timings from the source do not apply to a cut version of it, so the transcript has to be retimed onto the new edit before any phrase can be located reliably.

Step six is narrower than it sounds. A figure is often several words — a currency, a number, sometimes a unit — and colouring only the numeral leaves the rest plain, so Vizard Agent treats the whole spoken phrase as the thing to highlight.

Steps eight and nine are separate on purpose. Checking "the highlights" as a group is how one of them gets missed, so Vizard Agent looks at each one on its own frame at full size.

What does the result look like?

Captions that read as plain text with two coloured phrases in them, both landing exactly on the moment the speaker says the figure. Because nothing else in the video is coloured, the eye goes straight to them, which is precisely the effect that colouring every word destroys.

Vizard Agent shows you each highlight frame, so you can confirm the colour covers the whole amount rather than clipping the currency.

When does this not work well?

Emphasis needs scarcity. Two highlights in a thirty-second video work; eight do not, and at that point you are back to a preset that animates everything. Highlighting also does nothing for a figure the speaker garbles or that the transcript hears wrongly.

How do you fix a result that came back wrong?

Quote the phrase you meant, exactly as it is said. Because Vizard Agent matched against the spoken words in the first place, giving it the precise wording resolves a miss immediately, whereas describing the moment just sends it looking again in the same place.

How does Vizard Agent compare to doing it yourself?

By hand this means either accepting a preset that emphasises every word in turn, or editing individual caption events after the fact. The second is precise and tedious, and it breaks completely the moment the edit shifts by a second.

By hand Vizard Agent
Emphasis A preset, on every word Only the phrases you name
Locating the phrase Scrubbing From retimed word timings
What gets coloured The numeral The whole spoken phrase
Verification Watch it back Each highlight on its own frame
If the edit changes Redo by hand Retimed automatically

Common questions

Can it highlight words other than numbers? Yes. Name any phrase and Vizard Agent will find it.

How many highlights are too many? More than a few and the emphasis is lost. Vizard Agent will say so.

Will it colour the currency symbol? Yes, if you want the whole spoken phrase treated as one unit.

What if the transcript mishears it? Vizard Agent confirms the spoken text before applying colour.

Can I use more than one colour? You can, but Vizard Agent will point out that a single colour reads more clearly.

Does this work with grouped captions? Yes, and Vizard Agent recommends grouping over word-by-word here.

What if the phrase splits across screens? Vizard Agent adjusts the grouping so it stays together.

Will it check the placement? Yes. Vizard Agent checks it against faces and anything else in the lower frame.

Can it match my brand colour? Yes, subject to contrast against the footage.

Does the highlight animate? It can, though Vizard Agent finds a static recolour is usually stronger.

What if I re-cut the video? Vizard Agent retimes the transcript and the highlights follow.

Can it highlight in a translated track? Yes, though Vizard Agent will note the phrase may sit differently in the sentence.

Does it change the caption style? No. Vizard Agent changes the colour of the named phrases and nothing else.

Why does highlighting everything fail? Because emphasis is a comparison, and nothing is emphasised if everything is.