How to take the filler words out of your subtitles
List the filler words and where they appear, and tell Vizard Agent the picture must stay unbroken. Vizard Agent removes them from the captions at those exact timecodes, tracks down the brand typefaces the design needs, and rebuilds the subtitles without touching the footage or the audio.
What is the short version?
Automatic captions are faithful, which is the problem. They transcribe every "um" and every false start, and if the shot has to run unbroken you cannot fix that by cutting — the captions have to be corrected as text.
- Go to Vizard Agent and upload the reel with its captions.
- List the filler words and the timecodes where they appear.
- Say the person must stay on screen from start to finish.
What do you need before you start?
The video and a list of what to remove. Vizard Agent finds the words where you say they are, so timecodes make the correction exact rather than approximate. Naming the typefaces matters too, because rebuilt captions have to match the ones already there.
- The reel. With the captions as they stand.
- The words. The exact fillers to remove.
- The timecodes. Where each one appears.
- The typefaces. The brand fonts the captions use.
- The constraint. That the shot must not be cut.
What do you type into Vizard Agent?
Say what must not change alongside what should. Vizard Agent will otherwise reach for the obvious fix — cutting the filler out of the audio — which is usually right and is exactly what a continuous shot rules out.
Prompt
Variants worth knowing:
- Captions only. The audio and the picture untouched.
- A specific typeface. Matched rather than approximated.
- One reel of several. Corrected without redoing the set.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real subtitle correction job. Most of the visible effort goes somewhere entirely unexpected: hunting down the two brand typefaces, so that the rebuilt captions would match the untouched originals exactly.
- Checks the uploaded files and the project history.
- Extracts frames from the reference and the source to read the style and composition.
- Reviews the reference frames, then the frames from the reel itself.
- Checks which typefaces are available on the system.
- Searches specifically for the two brand fonts among the installed families.
- Downloads the first of them when it is not present locally.
- Searches code repositories for the second font's files.
- Searches the web for both, then for a download link for the second.
- Parses the download page, finds the direct link and retrieves the font — rather than substituting something that merely looks similar.
Steps 5 to 9 are the unglamorous half of any caption job. Rebuilt captions in a near-enough typeface look wrong beside the ones you did not touch, so finding the real font is what makes a correction invisible.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, HEVC, 60fps, 31.57 seconds, AAC audio. The same thirty-second reel, the same unbroken shot, the same audio — with the fillers gone from the captions and the type matching the rest of the set exactly.
Nothing was cut. The speaker still says "um" if you listen closely; the captions simply no longer write it down, which is what a viewer reading along actually notices.
When does this not work well?
Correcting captions without touching the audio deliberately opens a gap between what is heard and what is read, and that gap has limits. These are the points at which Vizard Agent will suggest cutting the audio after all, rather than pushing further.
- Captions and speech diverge. Remove too much and the reading stops matching.
- Heavy filler needs an audio cut. Past a point the speech itself is the problem.
- Brand fonts may be unfindable. Some are licensed and not downloadable.
- Burnt-in captions must be rebuilt. Text baked into the picture cannot be edited.
- Timecodes have to be right. A vague description costs a round.
How do you fix a result that came back wrong?
Point at the timecode where it is wrong. Vizard Agent keeps the caption timings, the located typefaces and all the frame samples, so a word can be removed or restored again without anybody having to rebuild the whole reel.
- "You missed one at 0:52." Removed from the caption at that timecode.
- "The font is not quite right." Re-rendered in the located brand typeface.
- "Now it does not match what she says." The word is restored.
How does Vizard Agent compare to doing it yourself?
By hand this means editing a caption file, re-rendering the subtitles, and discovering that your editor does not have the brand font — so the corrected lines are set in something almost right, which reads as worse than the fillers did.
| By hand | Vizard Agent | |
|---|---|---|
| The fix | Cut the audio, breaking the shot | Captions corrected, picture untouched |
| Typefaces | Substitute something similar | Hunted down and retrieved |
| Precision | Scrub for each filler | Removed at the stated timecodes |
| The rest of the set | Drifts from the corrected one | Matched to the originals |
Common questions
Does the audio change? No. Vizard Agent edits the captions only, leaving the speech and the picture as they are.
Why not just cut the "um" out? Because cutting the audio cuts the picture, and some shots have to run unbroken.
Do I need to give timecodes? They make it exact. Vizard Agent can search for them, but a list is faster and more reliable.
What if my brand font is not installed? Vizard Agent goes looking for it rather than substituting something that looks close.
Can it fix captions that are burnt into the video? Only by rebuilding them, since baked-in text is pixels rather than editable subtitles.
Will the captions still match the speech? Closely. Remove a few fillers and nobody notices; remove whole phrases and the reading drifts.
Can it correct one reel out of a set? Yes, and match it to the others. Vizard Agent reads the existing caption style off them.
Does it work in any language? Yes, provided a typeface covering the script is available.
What if the font is licensed? Then it may not be retrievable. Vizard Agent will say so rather than substituting silently.
Can it clean the captions on a whole batch? Yes. Vizard Agent applies the same word list across every reel in a set, which matters because a filler removed from one and left in another is more noticeable than either on its own.
Is it worth removing every single filler? No. Vizard Agent leaves the ones nobody notices, because captions stripped of every hesitation start reading like a script rather than like the person actually speaking.
Does this help accessibility? It helps readability rather than accuracy. Vizard Agent keeps the captions faithful to what is said while dropping the noise, which is the balance a viewer reading along actually wants.
Does Vizard Agent check the result? Yes. It reviews the rebuilt captions against the original style before delivery.