Vizard Agent

How to produce a second language version with the same cut

Last updated 2026-09-01 · 8 min read

Say up front that a second language version is coming. Vizard Agent keeps the original speech and the identical cut, translates the captions and every piece of on-screen text, and applies any later change to both versions together so the two never drift apart.

What is the short version?

The instinct is to re-dub, and for most videos that is wrong. If the speaker's own voice is the reason the video works — an interview, a testimonial, an appeal — you keep the audio and translate everything written on top of it.

  1. Go to Vizard Agent and build the first version normally.
  2. Say a second language version is coming, so the edit stays portable.
  3. Name the graphics that carry text: titles, figures, the end card.

What do you need before you start?

The source material, and an early decision about the voice. Vizard Agent keeps the original audio whenever you ask it to, and that decision matters most when the speaker's accent, hesitation and delivery are a large part of what makes the video credible in the first place.

What do you type into Vizard Agent?

Say the second version is coming before the first one has been built, not afterwards. Vizard Agent then keeps every piece of text separable from the picture underneath it, and that separation is exactly what makes the translated version a re-render rather than a rebuild from scratch.

Prompt

Make a [25–35] second [Spanish] short from this interview. Keep the original [Spanish] audio, natural pacing and original framing — no stretching, no automatic speaker cropping. Clean readable [Spanish] captions. Open with "[10.000 LITROS DE AGUA AL MES]". Show these figures one at a time: [...]. End for five seconds on "[MANTENGAMOS EL AGUA CORRIENDO]" and [the link]. Keep the edit easy to localise into a separate [English]-captioned version next.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real bilingual appeal. Steps twenty-seven onwards are the second version, and note what happens at the end: a change requested later is applied to both versions rather than to one.

  1. Looks at the interview and reads its details.
  2. Views the supporting photographs supplied as B-roll.
  3. Transcribes the interview for the exact words and their timings.
  4. Samples frames across it to see the framing.
  5. Grabs full-resolution frames to read the captions already burnt in.
  6. Listens closely to the opening words to check the transcript.
  7. Corrects the transcript into clean caption text rather than raw speech.
  8. Generates the styled captions and designs the titles, figure chips and end card.
  9. Renders the first cut, extracts frames across it and reviews them as a sheet.
  10. Fixes the end title size and the caption highlight colour and re-renders.
  11. Has the full render watched for sync and glitches.
  12. Builds the second language's captions and translated graphics.
  13. Renders the second version with identical cuts and timing.
  14. Reviews the second version's frames and uploads it.
  15. Applies a later change — a panel removed, a blur kept — to both versions and re-renders each.

Step fifteen is the one that saves the job. Bilingual sets drift because a note applied to one version is forgotten on the other, and Vizard Agent rebuilds both from the same edit rather than treating them as separate projects.

What does the result look like?

The appeal this page is written from exists as two vertical shorts of twenty-five to thirty-five seconds, one captioned in each language, carrying the speaker's original audio in both versions and with every title, figure chip and end card translated across into the second one.

The cuts are identical. That is deliberate: two versions that differ in pacing are two videos to maintain, whereas two versions that share an edit are one video with a second caption track and a second set of cards.

When does this not work well?

Keeping the original audio is the right call when the voice itself carries the credibility, and the wrong call when the second audience simply will not sit and read. Vizard Agent will tell you when the translated captions cannot keep pace with how fast the speech runs.

How do you fix a result that came back wrong?

Say which version and which element you mean, or simply say "both". Vizard Agent keeps the shared edit, every caption track and all the graphics, so a correction can be applied to one version alone or propagated across every version you have at once.

How does Vizard Agent compare to doing it yourself?

By hand, the second version is a duplicated project file that immediately starts diverging — one gets the client's note, the other does not, and three weeks later they are different videos. Vizard Agent rebuilds both from the same edit.

By hand Vizard Agent
The edit Duplicated, then drifts Shared by both versions
Graphics Re-made per language Re-rendered from the same design
A late change Applied to whichever you open Applied to every version
Captions Re-timed by hand Re-timed from the same word timings

Common questions

Should I dub or subtitle? Subtitle when the speaker's voice is the point. Vizard Agent keeps the original audio if you ask.

Does the on-screen text get translated too? Yes, and this is usually forgotten. Vizard Agent re-renders every card and figure.

Will the two versions match? Exactly. Vizard Agent builds them from the same cut and timing.

What if I change my mind later? Say "in both". Vizard Agent applies the change to every version and re-renders each.

Can I add a third language? Yes. Vizard Agent produces it from the same edit rather than starting again.

What about text already in the footage? It has to be handled. Vizard Agent blurs or covers burnt-in subtitles in the source.

Do figures need changing? Sometimes. Vizard Agent will convert currency or units if you say which.

Can the link differ per version? Yes. Vizard Agent swaps the end card per audience.

Will the captions fit? Usually. Vizard Agent flags lines that run long once translated.

Can it clone the speaker into the other language? Yes, if you want a dub. Vizard Agent can do that instead of captions.

Why keep the original voice at all? Because in an interview or an appeal, the accent, the hesitation and the room are the evidence. A polished dub makes it sound like an advertisement, which is precisely what you were avoiding.

Does it work for long videos? Yes. Vizard Agent re-times the captions from the same word timings whatever the length.

Does it check both versions? Yes. Vizard Agent reviews frames across each one before delivering.