Vizard Agent

How to land on-screen words exactly on the vocal that says them

Last updated 2026-09-04 · 7 min read

Tell Vizard Agent which words land on which line and to time them to the vocal rather than the beat. It measures where each phrase actually starts by analysing the signal, subtracts the animation's own pop-in delay, and checks on real frames that the word is readable at the moment it is sung.

What is the short version?

Text timed to a beat grid is close, and close is exactly what the eye catches. A chorus does not always enter on the downbeat, an animated word takes a few frames to become readable, and both errors push in the same direction — late.

  1. Go to Vizard Agent and send the track and the footage.
  2. Say which words go with which line of the vocal.
  3. Say they must land on the vocal, not on the beat.

What do you need before you start?

The master audio and the words themselves. Vizard Agent measures the onsets on its own, but it does need to know which phrase you mean — a word like "freedom" appearing four times across a song is four separate timing problems to solve, not one problem solved once.

What do you type into Vizard Agent?

Say "on the vocal" explicitly. Timing to the beat grid is the sensible default for a montage and the wrong default here, and the phrase that overrides it is worth saying plainly rather than implying with "make it tight".

Prompt

Put the on-screen words on the vocal that sings them, not on the beat grid. Measure where each phrase actually enters — the choruses do not all start in the same place — and account for the animation's own delay.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real music video whose on-screen text was landing consistently a fraction late. Steps four and five are two entirely separate causes of the same symptom, and Vizard Agent found and fixed them independently.

  1. Compares the current text times against the beat grid.
  2. Measures the exact vocal onsets in every chorus and chant section.
  3. Locates the onsets by signal analysis where the transcript was unclear.
  4. Tries word-level transcription on the chorus slices as a cross-check.
  5. Measures the pop-in delay of every text graphic.
  6. Re-times all the text hits and renders a new version.
  7. Verifies on frames that the text is up at each vocal hit.
  8. Checks when one word actually settles into readability.
  9. Views that settle strip frame by frame.
  10. Re-renders with the corrected timing for that word.
  11. Verifies it at all four of its appearances.
  12. Correlates the first chorus against the confirmed second one.
  13. Listens closely to where the first chorus really enters.
  14. Re-cuts it to the true vocal entry and re-renders.
  15. Checks the frames at each vocal moment in that section.

Step five is the one nobody expects. An animated word does not appear — it arrives, over several frames, and the moment it becomes readable is later than the moment it was triggered. If you time the trigger to the vocal, the word is legible after the syllable has gone.

Steps twelve to fourteen are the other half. Two choruses that sound identical are rarely identical, and assuming the first enters at the same offset as the second is how a video ends up right in one place and wrong in another.

What does the result look like?

Words that are already readable at the moment the line is sung, in every repeat, including the choruses that enter slightly early or late. Nothing Vizard Agent does here draws attention to itself, which is the entire measure of success for this kind of work.

The difference between this and beat-grid timing is a few frames. A few frames is also the difference between text that feels performed and text that feels stuck on afterwards.

When does this not work well?

Onset measurement needs a vocal that can actually be found in the signal. Where the voice sits deep inside a dense mix, or the phrasing is deliberately loose and behind the beat, the measurement gets less certain and the timing becomes a judgement call rather than a number.

How do you fix a result that came back wrong?

Name the word and say which occurrence of it you mean. Vizard Agent keeps the measured onsets, the per-graphic pop-in delays and the render itself, so nudging one appearance will not disturb the other three or force the whole timing pass to run again.

How does Vizard Agent compare to doing it yourself?

By hand this is done by ear and by nudging, one graphic at a time, and it is genuinely satisfying work when there are six words. At forty, the last twenty get less attention than the first twenty, and that is where the video stops feeling tight.

By hand Vizard Agent
Finding the vocal entry By ear, on a waveform Measured by signal analysis
Repeated phrases Assumed identical Each measured separately
Animation delay Rarely accounted for Measured per graphic
Verifying Play it back Checked on rendered frames
Forty words Diminishing attention The same treatment throughout

Common questions

Is the beat grid useless? No — Vizard Agent still builds one. It just is not what the words are timed to.

Do I need stems? No, but they help. Vizard Agent measures from the mix if that is all there is.

What about instrumental sections? Vizard Agent puts those on the beat. There is no vocal to land on.

Can it handle a rapped verse? Yes. Vizard Agent measures every utterance rather than sampling.

Will it catch a chorus that enters early? Yes — that is exactly what step twelve is for.

Does the animation style change the timing? Yes, and Vizard Agent measures each style's own delay.

Can I use my own graphics? Yes. Send them and Vizard Agent times them.

What if the words are wrong? Give the correct lyrics. Vizard Agent times what you supply.

Does it work for captions too? Yes, though speech is an easier case than singing.

Can it show me the frames? Yes. Vizard Agent verifies on frames and can hand them to you.

How tight is tight? Within a frame or two, which is below what the eye reads as late.

Does the audio ever move? No. Vizard Agent moves the graphics to the audio, never the reverse.

Can it do this for a live recording? Yes, and it matters more — live phrasing drifts from the grid constantly.

Why not just eyeball it? Because the error is a few frames in one direction, and a few frames is precisely the amount the eye notices and the ear cannot explain.