Vizard Agent

How to bleep a word and hide it in the captions at the same time

Last updated 2026-09-04 · 7 min read

Tell Vizard Agent to bleep the words and mask them in the captions together. It lists the sensitive words out of the transcript, covers them in both places, and then runs the bleeped audio back through a transcriber — if the word still comes out, the bleep did not land where it needed to.

What is the short version?

A bleep over the audio and a caption that still spells the word out is not a censored video, it is a video with a bleep in it. Both channels carry the word, and platforms and viewers read both.

  1. Go to Vizard Agent and say which words cannot stay in.
  2. Say both the audio and the captions have to be covered.
  3. Ask it to check the bleeped audio against a transcriber.

What do you need before you start?

The footage and a rule about what cannot stay in. Vizard Agent can produce the candidate list itself by reading the transcript, which is usually better than trying to remember every instance — a long recording contains words you have completely forgotten saying.

What do you type into Vizard Agent?

Ask for the list first if you are not sure. Handing over the judgement of what counts as sensitive is reasonable for a first pass, and reviewing a list is much faster than scrubbing a recording listening for problems.

Prompt

For the YouTube version, mask the words that could get the video blocked. Bleep them in the audio and cover them in the captions too. Show me the list of words you found before you do it.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real long-form video prepared for a platform with content rules. Step nine is the one nobody thinks to do and the only one that actually proves it worked.

  1. Transcribes the footage and reads the full transcript.
  2. Finds and fixes transcription errors before acting on the text.
  3. Checks one unclear word rather than guessing at it.
  4. Lists the sensitive words out of the transcript.
  5. Masks them in the captions and builds the covering graphics.
  6. Checks the masked words on rendered frames.
  7. Locates the exact word timings for the audio bleeps.
  8. Creates the bleeped audio and tests it against the picture.
  9. Tests whether a transcriber still understands the bleeped words.
  10. Adjusts the bleeps and re-checks the punch-ins around them.
  11. Renders the video with the bleeps and the masks together.
  12. Verifies both on the final render.
  13. Measures the bleep and the stability of the shots around it.

Step nine is the whole point of this page. A bleep placed from a word timing can miss by a few frames and leave the first or last syllable exposed — inaudible to you on the tenth listen, perfectly audible to a machine, and machines are what scan uploads. Running the bleeped audio back through a transcriber turns "I think that is covered" into a check.

Step two matters more than it looks. If the transcript misheard the word, the bleep goes in the wrong place, so the text is corrected before anything is built on it.

What does the result look like?

A version where the word is gone from both channels — silent under the bleep, covered on screen — and where the cover is part of the design rather than an obvious black box, because it was built as a graphic rather than applied as a patch.

You also get the list of what was masked. That list is the thing to keep, because the next episode will contain most of the same words and the judgement has already been made once.

When does this not work well?

Masking works on individual words. It does not work on meaning, and a video that is unmistakably about the thing you bleeped is not made safe by the bleep — Vizard Agent covers the word, not the subject the whole video is discussing.

How do you fix a result that came back wrong?

Say which instance. Vizard Agent keeps the word timings, the mask graphics and the verification result, so widening one bleep or covering a missed instance is a re-render of those seconds rather than a fresh pass over the recording.

How does Vizard Agent compare to doing it yourself?

By hand this is scrubbing a long recording listening for words, placing tones, then separately editing the caption file, then hoping. The hoping is the part that fails — bleeps that miss by four frames sound fine to the person who placed them.

By hand Vizard Agent
Finding the words Listening through Listed from the transcript
Placing the bleep By ear From word-level timings
The captions A separate edit Masked in the same pass
Proving it worked You hope Bleeped audio re-transcribed
The next episode Start again The word list carries over

Common questions

Can it decide what to mask? Yes. Ask Vizard Agent for the list and review it before it acts.

Will it bleep the wrong word? It works from corrected word timings. Vizard Agent fixes transcript errors first.

Can I use a sound other than a bleep? Yes. Vizard Agent will use a tone, a silence, or a sound effect you choose.

What covers the word on screen? Whatever you want. Vizard Agent can star it out or build a graphic.

Does it check its own work? Yes — that is step nine. The bleeped audio goes back through a transcriber.

Can I keep an uncensored version? Yes. Vizard Agent renders both from the same edit.

What about the burnt-in captions? Vizard Agent masks them in the same render, so the two never disagree.

Does it work in other languages? Yes. Vizard Agent works from the transcript in whatever language it is.

Can it mask a name instead? Yes. Vizard Agent masks any word or phrase you name.

Will the bleep sound abrupt? Vizard Agent measures it and adjusts rather than dropping a tone in.

Can I reuse the list next time? Yes. That is the most useful thing to keep.

Does it handle a word said twenty times? Yes. Vizard Agent covers every instance, each from its own timing.

What if two people talk at once? Vizard Agent tells you — the bleep will take both voices with it.

Why bother with the transcriber check? Because the failure here is silent. A bleep that misses by a few frames sounds fine to you and reads perfectly to whatever scans the upload.