Vizard Agent

How to make a short that teaches the difference between two similar-sounding words

Last updated 2026-09-05 · 8 min read

Give Vizard Agent the two words and the mistake learners make. It auditions a voice for the explanation and a separate one for the demonstration, tests how the explaining voice handles the foreign word inside its own sentences, and puts the phonetic spelling on screen next to each one.

What is the short version?

A pronunciation lesson has one requirement no other video has: the word being taught must be said correctly. Everything else — the pictures, the captions, the joke at the start — is packaging around a few seconds of audio that has to be right.

  1. Go to Vizard Agent with the two words and the confusion.
  2. Say which language explains and which one demonstrates.
  3. Ask it to check how the explaining voice says the foreign word.

What do you need before you start?

The pair of words and the mistake. The mistake is what makes it a video rather than a dictionary entry — a learner saying the wrong one in a real sentence is the hook, and it tells the lesson what it has to fix.

What do you type into Vizard Agent?

Say the two-voice structure explicitly. A single narrator reading a sentence that switches language for two words is the default, and it is the one thing that ruins this format — the demonstration comes out with the explaining language's accent.

Prompt

Make a vertical short teaching the difference between [walk] and [work] for [Mandarin] speakers. Open with a learner getting it wrong in a sentence. Explanation in [Mandarin], demonstration in English by a separate native voice. Put the phonetic spelling on screen for both, and end on a tongue-twister using both words.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real pronunciation short built for learners of one language by speakers of another. Step four is the single test that decides whether this format works at all, and Vizard Agent runs it before anything else gets built.

  1. Checks the effects library and the voice options.
  2. Checks the motion cards, the caption tools and the typefaces.
  3. Auditions a narrator voice in the explaining language.
  4. Tests how that voice pronounces the foreign words inside its sentences.
  5. Listens to the mixed reading to hear whether the demo word survives.
  6. Tests the typeface for the writing system and confirms it renders.
  7. Generates the situation pictures, the music and the effects.
  8. Generates the explanation and the demonstration narration separately.
  9. Tests the motion cards showing the words and their phonetic spelling.
  10. Places the demonstration audio against the cards.
  11. Builds the teaching cards, the opening and the closing.
  12. Rebuilds the tongue-twister card after reviewing it.
  13. Pulls word-level timings so the cards land on the words.
  14. Switches to a clearer static caption style after checking legibility.
  15. Checks the loudness difference between the music and the narration.
  16. Patches the end card to match its own colour and rounded corners.
  17. Re-records the whole narration in your own voice when you send a sample.
  18. Checks each explanation sentence is correct after the re-record.
  19. Rebuilds the timeline against the new narration and re-times the cards.
  20. Delivers a version with no voiceover alongside the finished one.

Steps four and five are the whole reason this article exists. A voice built for one language will read a foreign word with that language's sounds — which is fine in an ad and fatal here, because the mispronounced word is the thing being taught. Hearing it before committing is a two-minute check that saves the video.

Step eighteen is the other half. Once the narration is re-recorded in a different voice, every explanation sentence gets checked again — a pronunciation video that explains the rule incorrectly is worse than no video, and a re-record is exactly where that slips in.

What does the result look like?

A short that opens on the mistake, shows both words with their phonetic spelling, demonstrates each in a native voice, and closes on a phrase that uses both — with the explanation in the learner's own language throughout.

The two voices are the thing you notice without noticing. The lesson sounds like a teacher and a model speaker rather than one person doing an impression of both.

When does this not work well?

The format depends entirely on the difference being audible in a short clip played on a phone. Some contrasts need more than a demonstration to land, and some need a mouth rather than a voice — Vizard Agent will tell you when audio alone cannot carry the lesson.

How do you fix a result that came back wrong?

Say which word sounds wrong and which of the two voices said it. Vizard Agent keeps both narration tracks stored separately from the cards and the word timings, so re-recording one demonstration will not disturb the explanation track or the picture underneath it.

How does Vizard Agent compare to doing it yourself?

Teachers make these with a phone and their own voice, which is honest and works perfectly well. The difficulty is scale rather than quality: one pair of words per video, several videos a week, each of them needing clean audio, matching cards and phonetics that are actually correct.

By hand Vizard Agent
The demonstration Your own accent, or a friend's A separate native voice
Checking the mix Discovered on playback Tested before building
Phonetic spelling Typed and hoped Rendered and checked on screen
Your own voice The only option Optional, cloned from a sample
A second pair of words Start again The same structure, new words

Common questions

Why two voices? Because one voice reading both languages mispronounces the word being taught.

Can it use my voice? Yes. Send a sample and Vizard Agent re-records the explanation in it.

Will it get the phonetics right? Vizard Agent looks them up and shows them. Check them; you are the teacher.

Can it teach three words at once? Vizard Agent can, but one pair per short works far better.

What about tongue position? Say so and Vizard Agent adds a diagram; audio alone cannot show it.

Can it do a specific accent? Yes. Name it and Vizard Agent picks the demonstration voice for it.

Does the explanation have to be another language? No. Vizard Agent will do both in one language if that suits your audience.

Will the writing system render? Vizard Agent tests the typeface before building anything.

Can it add practice images? Yes. Vizard Agent builds pairs of pictures for each word.

What is the tongue-twister for? Practice using both words together. Vizard Agent builds it as the closing card.

Can I get a version without narration? Yes. Vizard Agent delivers a silent cut alongside the finished one.

Does it check the loudness? Yes. Vizard Agent measures the gap between the music and the voice.

How long should it be? Around a minute. Long enough for the rule, the demo and the practice.

Why check the mixed reading first? Because the demonstration is the lesson, and finding out it was wrong after the cards, the pictures and the music are built is an expensive way to learn it.