Vizard Agent

How to censor swearing in the captions without bleeping the audio

Last updated 2026-09-01 · 8 min read

Say you want the captions censored and the audio left alone. Vizard Agent stars out the words in the caption track only, so a viewer scrolling with the sound off sees nothing they would have to explain, while anyone actually listening hears the conversation as it happened.

What is the short version?

Bleeping the audio changes how the whole conversation sounds — it takes a relaxed exchange between two people and makes it feel policed. Starring the word out in the caption instead solves the problem you actually have, which is a screen full of profanity sitting in somebody's feed at work.

  1. Go to Vizard Agent and upload the recording.
  2. Say the captions should be censored and the audio should not.
  3. Say which words, or let it use the obvious ones.

What do you need before you start?

The recording and a decision about who this is for. Vizard Agent applies whatever list you give it, and the list is genuinely a judgement call — what a business podcast stars out is not what a comedy channel does.

What do you type into Vizard Agent?

Say clearly which of the two layers gets censored and which does not. Vizard Agent will otherwise assume you want both treated the same way, because a bleep normally accompanies the asterisks — and the entire point of this job is that the two come apart.

Prompt

Cut [5] vertical clips from this podcast. Kinetic captions in the middle of the frame. Censor spoken profanity in the caption text with asterisks, but **leave the source audio intact** — do not bleep anything. Must include the section about [X] and the section about [Y].

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real set of clips cut from a fifty-minute conversation. The censoring is a small part of a larger job, and it happens in the caption generation rather than anywhere near the audio.

  1. Checks the project files and the client's previous renders for the house style.
  2. Transcribes the full recording with word-level timings.
  3. Finds the mandatory sections the brief named.
  4. Fingerprints the camera setups used across the recording.
  5. Calibrates a centre crop per speaker for the single-speaker shots.
  6. Uses speaker dominance to crop the wide two-shots onto whoever is talking.
  7. Adjusts the wide-shot centring to keep the headroom balanced.
  8. Builds the vertical layout — a header band for the headline, the video block beneath.
  9. Generates the kinetic captions positioned mid-screen.
  10. Applies the profanity list to the caption text only, replacing the letters with asterisks.
  11. Leaves the speech audio untouched through the whole pipeline.
  12. Normalises each clip's speech to a broadcast loudness target with smooth fades.
  13. Renders the clips and the separate promotional cut.
  14. Reviews the frames for framing, captions and the header band.

Step ten and step eleven are deliberately separate operations. The censoring happens to the text that gets drawn on screen, and nothing in the audio chain ever knows about it.

What does the result look like?

The clips this page is written from were vertical at 1080x1920, built with a 250-pixel header band carrying the headline above the video block, kinetic captions positioned mid-screen, and every clip's speech normalised to a broadcast loudness target with smooth fades at each end.

On screen the swearing appears as asterisks. In the audio it is exactly as it was said. Nobody scrolling past has anything to explain, and nobody who chooses to listen gets a bleep in their ear.

When does this not work well?

Censoring a caption is fundamentally a judgement about your own audience, and there is no setting anywhere that will make that judgement on your behalf. Vizard Agent applies whatever list you supply precisely, which means the list itself is the entire decision.

How do you fix a result that came back wrong?

Adjust the word list rather than going back into the clip itself. Vizard Agent regenerates the whole caption track from the same transcript, so adding or removing a single word is a re-render of the text layer rather than a new edit of the video underneath.

How does Vizard Agent compare to doing it yourself?

By hand, censoring means either editing a caption file by hand for every clip, or reaching for a bleep and changing how the whole conversation feels. Vizard Agent treats the two layers as separate decisions, because they are.

By hand Vizard Agent
The caption Edited word by word Generated from a list you set
The audio Bleeped, usually Untouched unless you ask
Across clips Done again per clip Applied to the whole set at once
A changed list Every file re-edited One re-render of the text layer

Common questions

Why not just bleep it? Because a bleep changes the tone of the conversation. Starring the caption fixes the screen without touching the sound.

Will people still hear the word? Yes, if they have the sound on. This is not a clean version of the audio.

Who decides the list? You do. Vizard Agent applies exactly what you supply.

Can it keep the first letter? Yes. Vizard Agent will star the rest and leave it recognisable.

Does it work across a whole batch? Yes. Vizard Agent applies the same list to every clip in the set.

What if a word is mistranscribed? It can be starred wrongly. Vizard Agent shows the transcript so you can check.

Can it bleep the audio too? Yes, as a separate pass. Vizard Agent keeps the two operations independent.

Does it affect the loudness? No. Vizard Agent normalises the speech separately from anything the captions do.

Can I censor other words? Yes — brand names, a competitor, a person's name. Vizard Agent takes any list.

Will the captions still be readable? Yes. Vizard Agent keeps the asterisks in the same style as the surrounding text.

Is this dishonest? No, and the direction matters. Nothing is being added or claimed — the audio is exactly as recorded, and the screen is simply made safer to be seen in public.

Does it work on burnt-in captions I already have? Those would have to be removed and rebuilt. Vizard Agent can do it, but generating fresh captions is cleaner.

Does it check the result? Yes. Vizard Agent reviews the rendered frames to confirm the captions read correctly and that the starring has landed on the right words.

Can it censor a name instead? Yes. Vizard Agent takes any list of words, including a person or a competitor you would rather not name on screen.