Vizard Agent

How to subtitle a video that switches language mid-conversation

Last updated 2026-08-30 · 7 min read

Upload the footage and give Vizard Agent the rule: which language gets subtitled into which. Vizard Agent transcribes everything, finds where each language starts and stops, reads the structure of the bilingual passages, and subtitles each stretch in the language its own speaker was not using.

What is the short version?

Footage where people switch language partway through is unusable to half of each audience. Subtitling it properly means the subtitle language has to change whenever the spoken language does, which is a rule Vizard Agent applies per passage rather than a single setting.

  1. Go to Vizard Agent and upload the footage.
  2. State the rule — English subtitles on the Arabic, Arabic on the English.
  3. Say how many clips and how long.

What do you need before you start?

The footage and the rule. Vizard Agent detects where each language is spoken and locates the boundaries itself, so you do not have to mark the switches. What you do have to decide is which audience the subtitles serve at each moment.

What do you type into Vizard Agent?

State the rule as a pair. Vizard Agent applies it per passage rather than per file, so "English subtitles on the Arabic speech and Arabic subtitles on the English speech" is a complete instruction that does not need any timestamps from you.

Prompt

Cut [3] vertical reels from this mixed [Arabic and English] footage. Put [English] subtitles on the [Arabic] speech and [Arabic] subtitles on the [English] speech. Add a [YouTube] call to action at the end.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real bilingual repurposing job. The step that makes the rule workable is the eleventh: it searched the transcript specifically for the Arabic passages rather than treating the recording as one language.

  1. Checks the project history and files, then the download options.
  2. Downloads the source video, retrying when the first attempt fails.
  3. Inspects the downloaded file and transcribes and surveys the footage.
  4. Looks at frames across the video and reads the visual shot log.
  5. Checks the transcript's structure and saves it.
  6. Searches the transcript for the passages in the second language.
  7. Reads the transcript through, including the end.
  8. Loads the clip-repurposing playbook and checks the caption and effect options.
  9. Checks the auto-reframe tool, examines the candidate moments and the bilingual conversation specifically, reads its breakdown, then checks the framing of each candidate.

Step 6 is the one that separates this from ordinary captioning. A transcript of mixed speech is not automatically labelled by language, so finding where the switches fall is a deliberate search rather than something that falls out of the transcription.

What does the result look like?

A set of vertical reels cut from the bilingual footage, each one subtitled so that whenever the speech changes language the subtitle language changes with it, with a call to action at the end and framing checked at every chosen moment.

Both audiences can follow the whole clip. That is the entire objective, and it is why the rule is stated as a pair rather than as a single subtitle language applied throughout.

When does this not work well?

Mixed-language speech is considerably harder to transcribe accurately than either language alone, and some switching patterns defeat any subtitle rule Vizard Agent could reasonably apply. These are the limits worth knowing before you plan a channel around the format.

How do you fix a result that came back wrong?

Point at the passage that is wrong and say what it should read. Vizard Agent keeps the transcript, the located language boundaries, the shot log and the framing checks, so one stretch can be re-subtitled without redoing the transcription.

How does Vizard Agent compare to doing it yourself?

By hand this means transcribing two languages, marking every switch, translating each passage into the other language, and then styling two subtitle tracks that never appear at the same time. It is the reason most bilingual footage never gets repurposed at all.

By hand Vizard Agent
Finding switches Listen and mark Searched for in the transcript
Translation Passage by passage Applied per passage, by rule
Typefaces Discover the boxes late Checked for the scripts involved
Framing Crop and hope Checked at every candidate moment

Common questions

Do I have to mark where the language changes? No. Vizard Agent searches the transcript for each language and locates the boundaries.

Can both subtitle languages appear at once? They can, but the frame gets crowded. Vizard Agent finds alternating reads better.

Will right-to-left text display correctly? Yes, provided a supporting typeface exists. Vizard Agent checks before rendering.

How accurate is the transcription? Lower than for a single language. Code-switching is genuinely hard, so review the result.

Can it add a call to action? Yes. Vizard Agent puts it in whichever language the clip ends in.

Why state the rule as a pair? Because the point is that each audience can follow the whole clip, and that only works if the subtitle language is whichever one the speaker is not currently using.

What if someone switches mid-sentence? That is the hard case. Vizard Agent handles passages cleanly and will flag the messy ones.

Can I get one clip per language instead? Yes, though that loses the conversation, which is often the interesting part.

How many reels from one recording? As many as the footage supports. Vizard Agent checks the framing of each candidate.

Does this work for any language pair? Wherever both are transcribable and the scripts have typefaces. Vizard Agent checks both before rendering, since a missing font for one side of the pair fails silently until the render.

Should I subtitle the whole thing in both languages instead? It is an option, and it costs frame space. Vizard Agent will build it, but two permanent subtitle tracks leave very little room for anything else in a vertical frame.

Why not just dub one language into the other? Because the switching is often the point. Vizard Agent can dub if you prefer, but a conversation that moves between languages loses something real when it is flattened into one.

How many languages can be mixed? Two is the common case and the one Vizard Agent handles most cleanly. Three becomes hard to read, because the subtitle rule needs a target for every speaker in every passage.

Does Vizard Agent check the finished reels? Yes. It reviews the framing and the subtitle timing before delivery.