How to caption a video with your exact words and the audio's timings
Give Vizard Agent the script or lyric file and say the captions must use that wording exactly. It builds its own word-level transcript purely to find where each phrase is sung or spoken, then lays your text onto those spans — including in passages the recogniser could not hear, where your lines are placed at the actual vocal entrances.
What is the short version?
When you have the words already, transcription is the wrong tool for the job. It will get a name wrong, tidy a lyric, or drop a quiet line — and every one of those is a mistake you could have avoided by not asking the question.
- Give Vizard Agent the exact wording as a file.
- Say the timings should come from the audio.
- Point out any quiet passages in advance.
What do you need before you start?
The audio and the words, in the order they occur. Vizard Agent aligns your text against the audio phrase by phrase, so the file needs to match what is actually performed — a lyric sheet missing a repeated chorus will throw the alignment off from that point on.
- The audio or video. The performance itself.
- The wording. A script, SRT or lyric file.
- In performance order. Including repeats.
- The quiet sections. Outros, whispers, backing lines.
- The caption look. So the alignment renders into it.
What do you type into Vizard Agent?
Separate the two sources explicitly. The instruction that gets this right names where the text comes from and where the timing comes from as two different things, because the default behaviour is to take both of them from the same place — the recogniser — and that is precisely the arrangement you are trying to break.
Prompt
Variants worth knowing:
- "Use this text word for word." Your file is the authority.
- "Take the timings from the audio." The other half of the job.
- "Including the quiet outro." Where lines go missing.
What does Vizard Agent actually do?
Here is the Vizard Agent sequence on a real music video where a supplied lyric sheet had to be married to the vocal performance. The recogniser is still used throughout — but only ever as a clock, never as a source of words.
- Inspects the current captions for drift against the vocal.
- Builds an independent word-timed transcript to compare against.
- Compares the opening captions with the first vocal entrances.
- Checks the later sections for drift that has accumulated.
- Reads the word-timing structure to find the vocal spans.
- Aligns your supplied text onto those spans.
- Renders the aligned captions into the existing design.
- Isolates the quiet outro and analyses its sung phrase timing.
- Places the supplied lines at the real entrances the recogniser missed.
- Builds the combined track and re-checks it end to end.
Step two is counter-intuitive and important. The transcript is generated and then never used for its words — it exists only to say where in the audio each syllable falls, which is the one thing your script cannot tell anyone.
Step eight is where most automatic captioning quietly fails. A quiet passage produces no recognised words, so the lines belonging to it simply disappear, and because the lines are quiet nobody notices they are gone.
Step nine is the repair. Vizard Agent analyses the phrase timing of the quiet section directly and puts your lines where the voice actually enters, rather than leaving a gap or stretching the previous caption across it.
What does the result look like?
Captions that say what you wrote, when it was said. Names are spelled your way, lyrics read as written, the quiet lines are present, and Vizard Agent tells you which sections it had to align by phrase rather than by word.
When does this not work well?
Alignment needs the text and the performance to agree. An improvised take, a dropped verse, or a different edit of the song will put your file and the audio out of step, and from that point the alignment is guessing.
- Improvised or changed lines. The file no longer matches.
- A different mix or edit. Verses in another order.
- Heavy overlapping vocals. No single line to align to.
- Instrumental passages. Nothing to attach a line to.
- Text that is a translation. Alignment needs the spoken language.
How do you fix a result that came back wrong?
Say where the drift starts rather than where you happened to notice it. Alignment errors propagate forward, so the useful piece of information is the first line that is wrong, and Vizard Agent re-aligns from that phrase onward instead of shifting the whole track.
- "It drifts from the second chorus." Re-aligned from that phrase.
- "The outro lines are missing." That section is analysed separately.
- "It changed my wording." The supplied file is re-imposed.
- "Two lines are merged." The phrase boundary is re-cut.
How does Vizard Agent compare to doing it yourself?
By hand you either type your words over an auto transcript, which is fast and re-times nothing, or you place each line by ear, which is accurate and takes an afternoon. The first gives you the right words at the wrong moments.
| By hand | Vizard Agent | |
|---|---|---|
| The wording | Retyped over a transcript | Taken from your file |
| The timing | Inherited or done by ear | Aligned to the vocal |
| Quiet passages | Lines quietly disappear | Analysed separately and restored |
| Drift | Accumulates unseen | Checked section by section |
| Fixing one line | Shifts the rest | Re-aligned from that phrase |
Common questions
Why transcribe at all if I have the words? For the timings. Vizard Agent uses the transcript as a clock, not as text.
What format should my file be? A script, an SRT or a plain lyric sheet. Vizard Agent reads all of them.
Will it keep my spelling? Yes. That is the point — Vizard Agent does not re-word your text.
What about a name the recogniser mangles? It never sees it. Your spelling is what Vizard Agent renders.
Can it handle repeated choruses? Yes, if the repeats are in your file. Vizard Agent matches them to the vocal.
What if a line is whispered? Vizard Agent analyses that section separately to find the entrance.
Does my caption design survive? Yes. Vizard Agent renders the new alignment into the existing look.
Can it do this for two languages at once? Yes, though Vizard Agent aligns to the spoken language and places the other alongside.
What if my file has an extra line? Vizard Agent flags it rather than forcing it into the audio.
How accurate is the alignment? To the word where the vocal is clear, to the phrase where it is not.
Can I get the aligned file back? Yes. Ask Vizard Agent for the corrected subtitle file as well as the video.
Does it work for speech as well as singing? Yes. The method is the same for a scripted voiceover.
What if the audio has an instrumental gap? Vizard Agent leaves it empty rather than holding the previous line across it.
Why did only the end drift? Usually a quiet passage. Tell Vizard Agent and that section gets its own analysis.