Vizard Agent

How to change a name in both the speech and the on-screen text

Last updated 2026-09-28 · 8 min read

Tell Vizard Agent the name has to change in the spoken audio and in the subtitles, and that everything else stays. It clones each speaker from their own clean sample, fits each new word into the interval the old one occupied at the same level, and rebuilds the subtitle at the same pixel position.

What is the short version?

A name is wrong and it sits in two places at once — said out loud and written on screen underneath. Fixing only one of them leaves the other one visibly contradicting it, which is worse than leaving both. Vizard Agent works both channels against the same single list of timings.

  1. Tell Vizard Agent to change it in the audio and the on-screen text.
  2. Let it clone each speaker who says the name.
  3. Check the frames where each subtitle opens and closes.

What do you need before you start?

The video and the exact replacements. Include inflected forms if the language has them — a name in a possessive or a question takes a different shape, and a list of one word per name will miss half the instances.

Say what must not change. "Keep the original video, background, music, timing, speaker voices and overall appearance" is the sentence that keeps this a patch instead of a re-edit, and it is worth writing out.

What do you type into Vizard Agent?

Name both channels explicitly, and give Vizard Agent whatever timestamps you already know. A brief that simply says "change this name" reliably gets one channel fixed and the other one left behind to contradict it, because nothing in the request said there were two.

Change this name to that one in both the spoken audio and the on-screen subtitles, around 0:00–0:03 and again at 0:10. Keep the voices, timing and everything else as they are.

Approximate timestamps are enough. They narrow where to look, and Vizard Agent finds the exact word boundaries from the transcript rather than from your estimate.

What does Vizard Agent actually do?

It builds one list and works both channels from it. A word-level transcript gives the precise start and end of every instance, and the frames at those moments get pulled at full resolution so the subtitle's appearance and position can be measured rather than guessed.

In that session the sequence ran:

Then the subtitles: their real pixel positions were measured so the new text sat in the same place, and a small visual trial confirmed the replacement blended with the background before anything was committed.

What does the result look like?

The same video with a different name in it, in both channels, agreeing with each other frame for frame. The voice is still the same speaker's own voice, the line runs to the same length as before, and the subtitle sits exactly where Vizard Agent found it.

Two rounds are normal. The first clones came from short samples and read stiffly, so longer, cleaner samples were used and the words regenerated, re-timed to the target durations and re-levelled — the second pass is where it stops sounding patched.

When does this not work well?

When the replacement is a noticeably different length from the original. A longer name still has to fit the gap the old one left behind, and past a small difference either the line runs over the words that follow it or the read has to be compressed to fit — Vizard Agent will tell you which of those two it is trading away rather than quietly picking one.

A name said with strong emotion is hard to match. A shout, a laugh mid-word or a name trailing into a sigh carries performance that a short cloned word does not, and those instances are the ones that stay noticeable.

And a name burnt into a moving graphic rather than sitting in a subtitle cue is a separate job again. The audio side of the work is unchanged, but the picture side stops being a text swap and becomes a tracked replacement that has to follow the graphic wherever it travels.

How do you fix a result that came back wrong?

If the old name flashes back briefly, look at the subtitle's opening and closing frames. That is the characteristic failure here — the replacement covers the middle of a cue and the original shows for a frame at each end, and the fix is each subtitle's in and out time rather than the text.

If the new word sounds stiff, ask for a longer voice sample. That single change did more in this session than any amount of re-timing, because a clone from a short sample has no natural cadence to borrow.

If the level jumps at the patch, say where. Vizard Agent measures the original name's loudness and matches the replacement to it rather than to the track average, and a mismatch means that measurement needs redoing on that instance.

Common questions

Does it change both channels by default? Only if you ask. Say audio and on-screen text explicitly.

How does it find every instance? Vizard Agent works from a word-level transcript, so the boundaries are exact rather than estimated.

What if two people say the name? Vizard Agent clones each speaker separately so each one says it themselves.

What about grammatical forms? List them for Vizard Agent, which generates each inflected version as its own word.

Can it match a question intonation? Yes. Vizard Agent generates a name said as a question as its own variant.

Will the line get longer? Not if the replacement fits. Vizard Agent trims the edges and re-times to the original interval.

Why does the old name flash back? The subtitle's boundary frames. Vizard Agent adjusts each cue's in and out times.

Does the music change? No. Vizard Agent touches only the replaced words.

Will the subtitle move? No. Vizard Agent measures its pixel position and puts the new text there.

How many passes should I expect? Two. Vizard Agent uses the first for timing and the second to make it sound natural.

Can it do several names at once? Yes. Vizard Agent works them all from one list of timings.

What if the name is in a graphic? For Vizard Agent that becomes a tracked replacement rather than a subtitle edit.

Does the speaker's voice survive? Yes, cloned from their own clean speech in the same recording.

Can I hear the patches before the render? Yes. Ask Vizard Agent for the generated words on their own, before anything is assembled.

Do I need to supply a voice sample? No. Vizard Agent takes a clean one from the speaker's own speech in the same recording.

What if one instance still sounds wrong? Name that instance. Vizard Agent regenerates only that word rather than the whole set.