Vizard Agent

How to make a count-the-word video where the number is actually right

Last updated 2026-09-10 · 8 min read

Ask Vizard Agent to index every occurrence of the word across the catalogue, then to check that index against the raw captions rather than trusting it. It reads each hit in context to discard the ones that mean something else, cuts the clips that survive, and counts what is actually in the finished video — which is the number you put on screen.

What is the short version?

A "count how many times we say it" video is a promise. If the number on the end card does not match what a viewer counts, the video fails at exactly the moment it was supposed to pay off.

  1. Go to Vizard Agent with the word and the catalogue.
  2. Ask for an index, cross-checked against the captions.
  3. Ask for the final count from the finished cut.

What do you need before you start?

The word and access to everything it might be in. Vizard Agent works from published transcripts where they exist, which means the catalogue can be searched before anything is downloaded — but the word has to be pinned down first, including which senses of it count.

What do you type into Vizard Agent?

Ask Vizard Agent for the index and the verification as a single instruction. A plain transcript search on its own gives you a list that is roughly right, and roughly right is the one thing this particular format cannot survive being.

Prompt

Make a vertical short where the viewer counts how many times we say this word. Index every occurrence across the whole catalogue, cross-check the index against the raw captions, read each one in context so we do not include the wrong sense, and tell me the exact count that ends up in the cut.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the sequence Vizard Agent worked through on a real vertical short built from eighty-six podcast episodes. Notice that the first several steps are all about counting accurately, and not one of them involves downloading a video file.

  1. Builds the episode list for the whole catalogue.
  2. Downloads the captions for every episode.
  3. Indexes every spoken occurrence of the word.
  4. Cross-checks that index against the raw captions.
  5. Reads every occurrence in context.
  6. Downloads only the sections that survived.
  7. Builds one audio strip of those sections.
  8. Transcribes the strip for exact word timings.
  9. Lists the precise in-clip position of every occurrence.
  10. Cuts each clip around its own word boundary.
  11. Checks every caption moment on the render.
  12. Counts what is in the finished cut for the answer.

Step four is the step that makes the number trustworthy. Published captions drop words, merge them, and mis-hear them, so an index built from a single pass is a starting point rather than an answer — Vizard Agent goes back to the raw text to check.

Step seven and step eight are the reason the cuts land. A catalogue timestamp tells you which episode and roughly where; it does not tell you which frame the word starts on, so Vizard Agent transcribes the assembled strip to get word boundaries inside the clips it actually has.

Step twelve is the discipline the format lives or dies on. The count is taken from the finished video rather than from the search that started it, because clips get dropped for length, framing or repetition along the way.

What does the result look like?

A short where every occurrence lands cleanly on the word itself, and an answer at the end that a viewer counting along at home will agree with. That agreement is the entire payoff of the format, and everything else is in service of it.

Vizard Agent also gives you the list of what made the cut and what did not, with reasons, so you can see exactly why the number is the number.

When does this not work well?

Some words cannot be counted reliably at all, and Vizard Agent will say so early. A word that appears inside other words, one that is frequently mumbled, or a catalogue with no usable transcripts behind it all make the index untrustworthy in ways that no amount of cross-checking will repair.

How do you fix a result that came back wrong?

Say which occurrence is wrong and in what way it is wrong. Because Vizard Agent keeps the index alongside the finished cut, adding or removing a single hit simply updates the count rather than forcing the whole catalogue to be searched again.

How does Vizard Agent compare to doing it yourself?

By hand this is a transcript search and a spreadsheet, and the number on the end card is the number the search returned rather than the number in the video. The two are almost never the same by the time the edit is finished.

By hand Vizard Agent
The index One transcript search Searched, then cross-checked
Wrong senses Caught by luck Read in context
Cut points Scrubbed by ear Word boundaries from the strip
The final number From the search From the finished cut
If a clip is dropped The number is now wrong The count is retaken

Common questions

Why check the index twice? Because published captions miss and mangle words. Vizard Agent verifies.

Do variants count? Your decision. Tell Vizard Agent and it applies it consistently.

What if there are too many? Vizard Agent trims to length and re-counts what remains.

Does it need the videos? Not at first. Transcripts are searched before anything downloads.

How does it place the cuts? From word timings taken off the assembled clips themselves.

Can it show a running counter? Yes, and Vizard Agent drives it from the same verified index.

What if the audio is unclear? Vizard Agent will say which occurrences it is unsure about.

Can it do a phrase instead? Yes, and phrases are usually cleaner to count than single words.

Will the number be on screen? If you want it. Vizard Agent takes it from the final cut.

Can it count across a whole channel? Yes. That is where the transcript-first approach pays.

What if an episode has no captions? Vizard Agent will transcribe it, or tell you it was excluded.

Does the order matter? Editorially yes. Vizard Agent will build to a peak if you ask.

Can I see what was excluded? Yes. Vizard Agent lists each exclusion with its reason.

Why do counts usually go wrong? Because the number comes from the search and the video comes from the edit.