Vizard Agent

How to make key words bigger inside your captions

Last updated 2026-09-22 · 8 min read

Ask Vizard Agent for captions with emphasised key words and say what should be emphasised. The styling is the easy half; the half that matters is that the captions are built from word-level timings, so the bigger word appears exactly as it is said rather than a beat late.

What is the short version?

Captions with one word larger and in colour are the standard look for short video now. They work when the emphasis lands on the word the line actually turns on, and they look arbitrary when it does not, which is why Vizard Agent needs a rule for choosing.

  1. Tell Vizard Agent to caption the video with emphasised key words.
  2. Say which words, or what kind of word, should be emphasised.
  3. Say where the captions must not sit.

What do you need before you start?

The video and a view on the emphasis. Vizard Agent can choose the words itself — the emotional ones, the numbers, the product names — but saying which kind you want is the difference between emphasis that reads as deliberate and emphasis that reads as decoration.

What do you type into Vizard Agent?

Say both parts of it: the look and the rule behind it. "Some key words larger and in a different colour" describes the look you want, while "the emotional ones" or "the numbers" is the rule Vizard Agent will apply when it decides which words get the treatment.

Prompt

Add captions to this video and make it look great. I want some key words in the captions to be a larger size than the rest and in a different colour — pick the emotional ones. Keep the captions clear of the main action, and make sure they are readable on a phone.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on a music-led vertical video captioned with emphasised words. The word timings get established first because every later decision depends on them, and the placement of the caption block is fixed afterwards once the emphasis exists.

  1. Inspects the project files and the uploaded video.
  2. Checks the video's duration and format.
  3. Checks which caption tools are available.
  4. Creates a visual sample and reviews the composition and the safe caption area.
  5. Transcribes the audio with word timings.
  6. Inspects the word-level timing file so the animation can be precise.
  7. Checks the styling options for emphasised keywords.
  8. Generates a vertical caption base with timed word reveals.
  9. Regenerates it using only the supported styling controls when the first attempt exceeds them.
  10. Applies larger, accent-coloured emphasis to the chosen keywords.
  11. Renders and reviews the captions across the edit, and measures the audio level.
  12. Lowers the caption block and re-renders so it clears the main action.
  13. Confirms the lowered captions stay legible and in sync.

Step five is the foundation. Line-level captions can be styled but not emphasised — a bigger word inside a line that appears all at once has no moment to land on, so word timings are what make the effect work at all.

Step nine is the practical reality of caption engines. Every one of them supports a particular set of controls, and a style built outside that set silently loses the part it cannot render.

Step twelve is the adjustment nearly every one of these needs. Emphasised words are taller than the rest, so a caption block positioned for even text starts covering the thing you are filming.

What does the result look like?

Captions that appear word by word with the speech, in which the chosen words arrive larger and in your accent colour on the beat they are said, sitting low enough to leave the action clear and readable at phone size.

When does this not work well?

Some material resists emphasis altogether. Fast speech leaves no room for a bigger word to register, dense captions become a mess of competing sizes, and a video in which Vizard Agent emphasises every word is a video with no emphasis in it at all.

How do you fix a result that came back wrong?

Say whether the problem is which words were chosen or how they look on screen. Vizard Agent changes the selection rule and the styling separately, so you are not re-rendering an entire caption treatment in order to change one accent colour.

How does Vizard Agent compare to doing it yourself?

By hand this is the most tedious caption job there is: split the line, size one word, colour it, time it, repeat for three minutes of speech. Most people do it for the first twenty seconds and let the rest run plain.

By hand Vizard Agent
Timing Line by line Word-level, from the transcript
Emphasis Chosen as you go A stated rule, applied throughout
Styling Whatever the editor offers Built within the engine's real controls
Placement Set once Adjusted for the taller emphasised words
Consistency Drifts after a minute The same throughout

Common questions

Can it pick the words itself? Yes. Give Vizard Agent a rule — emotional words, numbers, names — and it applies it.

Can I list the words instead? Yes, and Vizard Agent emphasises exactly those and nothing else.

How many words should be emphasised? One or two per line. More than that and Vizard Agent will tell you it stops reading.

Can the emphasis animate? Yes — a pop or a scale on the word as it lands.

Will it work in another language? Yes, and Vizard Agent checks the typeface supports its characters.

Can it match my usual caption style? Yes. Give Vizard Agent a previous video and it matches the treatment.

What about lyrics rather than speech? The same, and music videos are where this look is strongest.

Can it use two accent colours? Yes, though Vizard Agent will warn when it starts to look busy.

Will the captions be readable on a phone? That is checked on rendered frames rather than in a preview.

Can it put the captions at the top? Yes. Say where, and Vizard Agent keeps them clear of the action.

Does emphasis affect the timing? No. Vizard Agent keeps every word on its own moment.

Can I get the caption file? Yes, though the emphasis styling belongs to the rendered version.

What if the transcript is wrong? Correct it and Vizard Agent rebuilds the captions from the corrected text.

How do I check it? Watch muted at phone size: the emphasised words should tell the story alone.