How to make key words bigger inside your captions
Ask Vizard Agent for captions with emphasised key words and say what should be emphasised. The styling is the easy half; the half that matters is that the captions are built from word-level timings, so the bigger word appears exactly as it is said rather than a beat late.
What is the short version?
Captions with one word larger and in colour are the standard look for short video now. They work when the emphasis lands on the word the line actually turns on, and they look arbitrary when it does not, which is why Vizard Agent needs a rule for choosing.
- Tell Vizard Agent to caption the video with emphasised key words.
- Say which words, or what kind of word, should be emphasised.
- Say where the captions must not sit.
What do you need before you start?
The video and a view on the emphasis. Vizard Agent can choose the words itself — the emotional ones, the numbers, the product names — but saying which kind you want is the difference between emphasis that reads as deliberate and emphasis that reads as decoration.
- The video. With clear speech or lyrics.
- Which words. Named, or a rule for choosing them.
- The colours. Base and accent.
- Where they sit. Clear of faces and the action.
- The platform. It sets the safe area.
What do you type into Vizard Agent?
Say both parts of it: the look and the rule behind it. "Some key words larger and in a different colour" describes the look you want, while "the emotional ones" or "the numbers" is the rule Vizard Agent will apply when it decides which words get the treatment.
Prompt
Variants worth knowing:
- "Pick the emotional ones." The rule for choosing.
- "A larger size and a different colour." Both parts of the treatment.
- "Clear of the main action." Where they must not sit.
What does Vizard Agent actually do?
Here is the Vizard Agent sequence on a music-led vertical video captioned with emphasised words. The word timings get established first because every later decision depends on them, and the placement of the caption block is fixed afterwards once the emphasis exists.
- Inspects the project files and the uploaded video.
- Checks the video's duration and format.
- Checks which caption tools are available.
- Creates a visual sample and reviews the composition and the safe caption area.
- Transcribes the audio with word timings.
- Inspects the word-level timing file so the animation can be precise.
- Checks the styling options for emphasised keywords.
- Generates a vertical caption base with timed word reveals.
- Regenerates it using only the supported styling controls when the first attempt exceeds them.
- Applies larger, accent-coloured emphasis to the chosen keywords.
- Renders and reviews the captions across the edit, and measures the audio level.
- Lowers the caption block and re-renders so it clears the main action.
- Confirms the lowered captions stay legible and in sync.
Step five is the foundation. Line-level captions can be styled but not emphasised — a bigger word inside a line that appears all at once has no moment to land on, so word timings are what make the effect work at all.
Step nine is the practical reality of caption engines. Every one of them supports a particular set of controls, and a style built outside that set silently loses the part it cannot render.
Step twelve is the adjustment nearly every one of these needs. Emphasised words are taller than the rest, so a caption block positioned for even text starts covering the thing you are filming.
What does the result look like?
Captions that appear word by word with the speech, in which the chosen words arrive larger and in your accent colour on the beat they are said, sitting low enough to leave the action clear and readable at phone size.
When does this not work well?
Some material resists emphasis altogether. Fast speech leaves no room for a bigger word to register, dense captions become a mess of competing sizes, and a video in which Vizard Agent emphasises every word is a video with no emphasis in it at all.
- Very fast speech. The big word is gone before it reads.
- Too many emphasised words. Nothing stands out.
- Long lines. Size changes break the rhythm of reading.
- Busy backgrounds. The accent colour disappears.
- Poor audio. Word timings need clear speech.
How do you fix a result that came back wrong?
Say whether the problem is which words were chosen or how they look on screen. Vizard Agent changes the selection rule and the styling separately, so you are not re-rendering an entire caption treatment in order to change one accent colour.
- "The wrong words are emphasised." The selection rule is changed and reapplied.
- "The colour is hard to read." The accent is swapped and checked on the frames.
- "They cover her face." The block is lowered and re-rendered.
- "The timing is late." The captions are rebuilt from the word-level timings.
How does Vizard Agent compare to doing it yourself?
By hand this is the most tedious caption job there is: split the line, size one word, colour it, time it, repeat for three minutes of speech. Most people do it for the first twenty seconds and let the rest run plain.
| By hand | Vizard Agent | |
|---|---|---|
| Timing | Line by line | Word-level, from the transcript |
| Emphasis | Chosen as you go | A stated rule, applied throughout |
| Styling | Whatever the editor offers | Built within the engine's real controls |
| Placement | Set once | Adjusted for the taller emphasised words |
| Consistency | Drifts after a minute | The same throughout |
Common questions
Can it pick the words itself? Yes. Give Vizard Agent a rule — emotional words, numbers, names — and it applies it.
Can I list the words instead? Yes, and Vizard Agent emphasises exactly those and nothing else.
How many words should be emphasised? One or two per line. More than that and Vizard Agent will tell you it stops reading.
Can the emphasis animate? Yes — a pop or a scale on the word as it lands.
Will it work in another language? Yes, and Vizard Agent checks the typeface supports its characters.
Can it match my usual caption style? Yes. Give Vizard Agent a previous video and it matches the treatment.
What about lyrics rather than speech? The same, and music videos are where this look is strongest.
Can it use two accent colours? Yes, though Vizard Agent will warn when it starts to look busy.
Will the captions be readable on a phone? That is checked on rendered frames rather than in a preview.
Can it put the captions at the top? Yes. Say where, and Vizard Agent keeps them clear of the action.
Does emphasis affect the timing? No. Vizard Agent keeps every word on its own moment.
Can I get the caption file? Yes, though the emphasis styling belongs to the rendered version.
What if the transcript is wrong? Correct it and Vizard Agent rebuilds the captions from the corrected text.
How do I check it? Watch muted at phone size: the emphasised words should tell the story alone.