How to add animated verse callouts to a long lecture video
Upload the lecture and ask Vizard Agent to subtitle it and mark the passages the speaker quotes. Vizard Agent transcribes the whole talk, searches the transcript for references, looks up the real text of each one, and animates a card at the exact moment the quote begins.
What is the short version?
A recorded lecture is subtitled easily and annotated with difficulty. Finding every passage a speaker quotes across forty-four minutes means reading the whole transcript, matching each reference to the actual text, and placing a card on the word where the quote starts.
- Go to Vizard Agent and upload the lecture.
- Ask for professional subtitles and animated callouts for every quoted passage.
- Say which translation or wording the cards should use.
What do you need before you start?
The recording and a note about the wording you want. Vizard Agent transcribes, searches and looks up the passages itself, so the whole job runs from the file, and naming your preferred translation is the one detail that saves a revision round.
- The recording. Any length; this run was forty-four minutes.
- The translation. Cards quote whatever version you name.
- Subtitles too. Almost always, and they need to coexist with the cards.
- A card style. Or let Vizard Agent design and show you one.
- The delivery shape. Widescreen for a lecture; ask for clips separately.
What do you type into Vizard Agent?
Ask for both jobs in one instruction. Vizard Agent has to coordinate them — a card and a subtitle occupying the same part of the frame at the same moment is the failure case — and requesting them together is what lets it plan the gaps.
Prompt
Variants worth knowing:
- A specific translation. Name it and Vizard Agent uses that wording.
- Short clips as well. Each quoted passage makes a natural standalone clip.
- A card position. Say if the lower third has to stay clear.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real lecture. The interesting stretch is at the end: it rendered the whole video, checked the cards, found a timing error and re-rendered every segment rather than shipping it.
- Probes the file and grabs sample frames to see the framing.
- Transcribes the full lecture and reads both halves of the transcript.
- Searches the transcript for every reference the speaker makes.
- Looks up the exact text of each passage rather than transcribing what was said.
- Reads the word-level timings so each card lands on the word the quote begins on.
- Designs an animated card, composites it onto a real frame and diagnoses why the first render came out blank.
- Generates the subtitle file with gaps at the verses and verifies no subtitle appears under a card.
- Tests three subtitle styles — plain, higher-contrast white, and a dark backing box — and compares them on frames.
- Splits the video, renders it in parallel segments, finds the cards landed at the wrong timestamps, corrects the timing, re-renders every segment and verifies all eleven cards plus the subtitle resumption after each one.
Step 9 is the reason this article exists. Rendering a forty-four minute video once is slow; rendering it twice because you checked is what separates a finished job from a delivered one.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1920x1080, H.264, 30fps, 2631.73 seconds, AAC audio. Widescreen, forty-three minutes and fifty-two seconds, fully subtitled, with eleven animated verse cards placed on the words where the quotes begin.
Delivering the full lecture rather than a highlights cut is the right answer for this format. The video exists to be studied, and Vizard Agent annotated it in place rather than turning it into something shorter and less useful.
When does this not work well?
Vizard Agent works from what the speaker actually said, and quoting is a looser activity than it sounds. A paraphrase, a half-quote or a reference made without naming it are all harder to catch than a clean citation.
- Paraphrases are hard to detect. Vizard Agent finds explicit references reliably; allusions less so.
- Translations differ. Name the version or the card may not match the speaker's wording.
- Mishearings happen. Book names and numbers are exactly where transcription errors cluster.
- Long renders take time. Forty-four minutes at full quality is a substantial job, twice over if it needs correcting.
- Cards cover the picture. If the speaker is showing something, say where the card must not sit.
How do you fix a result that came back wrong?
Name the timestamp. Vizard Agent keeps the transcript, the word timings, the looked-up passage texts and every rendered card, so correcting one quote or moving one card is a re-render of that segment rather than the whole lecture.
- "That reference is wrong." Corrected and the card re-rendered.
- "He also quotes at 22 minutes." Added from the transcript Vizard Agent already has.
- "The subtitles clash with the card." The gap is widened and verified on frames.
How does Vizard Agent compare to doing it yourself?
By hand this is watching forty-four minutes with a notepad beside you, typing out eleven passages from memory, building eleven separate cards in a design tool, and then hand-suppressing the subtitle track underneath every single one of them.
| By hand | Vizard Agent | |
|---|---|---|
| Finding the quotes | Watch and note | Whole transcript searched |
| The verse text | Type it from memory | Looked up, not transcribed |
| Card placement | Eyeball the moment | On the word, from the timings |
| Subtitles under cards | Delete them manually | Gaps generated and verified |
Common questions
Is this only for scripture? No. Any video where a speaker quotes a fixed source works the same way — case law, poetry, a standards document, a company handbook. Vizard Agent searches the transcript for the references and looks up the real wording rather than trusting what the transcription heard.
How long a video can it handle? This run was forty-four minutes, rendered in parallel segments and delivered in full.
Does it use the real text of the passage? Yes. Vizard Agent looks each one up rather than relying on what the transcript heard.
Will the subtitles overlap the cards? No. Vizard Agent builds gaps into the subtitle file and checks frames to confirm.
Can it use my preferred translation? Yes. Name it in your instruction.
Does it render long videos in one pass? No. Vizard Agent splits the file and renders the segments in parallel, then joins them and re-adds the original audio, which is what makes a forty-four-minute job practical.
Can I get short clips of each quote too? Yes. Each passage is already located, so the clips come from the same analysis.
What if it misses one? Tell Vizard Agent roughly when, and it adds it from the transcript it already holds.
Does it check the finished render? Yes. Vizard Agent pulled verification frames from every card on this run and re-rendered all the segments when the timing was off.
Can it style the subtitles differently? Yes. Vizard Agent tested three styles on real frames here — plain, high-contrast white and a dark backing box.