How to add right-to-left Farsi captions to a talking-head reel
Upload the take to Vizard Agent and say the captions should be in Persian. Vizard Agent transcribes the speech, fetches a Persian font and the text-shaping libraries the script needs, renders one test caption and looks at it to confirm the letters join and run right to left, then cuts the reel around the verified result.
What is the short version?
Captions in a right-to-left script are not a settings change. The letters have to join to each other, the line has to run the other way, and most rendering pipelines silently produce disconnected letters in the wrong order without throwing a single error.
- Go to Vizard Agent and upload the talking-head take.
- Say what language the speech is in and what language the captions should be.
- Say the feed, the length and whether you want B-roll over the talk.
What do you need before you start?
The take and the language. Vizard Agent handles the font, the shaping and the direction itself, and the thing worth telling it explicitly is the script rather than only the language, since a language can be written more than one way and the rendering path differs.
- The take. One continuous recording is the normal input for this.
- The spoken language. Vizard Agent transcribes in it rather than translating first.
- The caption language and script. Persian, Arabic, Hebrew, Urdu. Say which.
- The topic. It steers what B-roll gets sourced to cover the talk.
- The feed. Vertical with a punch-in on the speaker is the convention here.
What do you type into Vizard Agent?
Say the language twice: once for the speech and once for the captions. Vizard Agent works out the font and the shaping from that, and being explicit about the script is what stops it treating a right-to-left language as an ordinary left-to-right one.
Prompt
Variants worth knowing:
- Captions in a second language. Ask for a translated track as well and Vizard Agent builds both.
- No B-roll. A tight punch-in cut on the speaker alone is a different and often stronger edit.
- Emoji in the captions. Say so; the font has to support them and Vizard Agent checks.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real Persian reel. The whole first third of the run is spent on making the script render correctly, and not one step of it would have been necessary if the captions were in English.
- Probes the uploaded video and reads the tool documentation.
- Transcribes the Persian audio rather than translating it first.
- Samples frames from the footage and looks at them to see the framing.
- Checks which Persian fonts and emoji are available on the system.
- Lists the installed font families and finds nothing suitable.
- Fetches a Persian font and installs the text-shaping libraries the script requires.
- Creates a bold instance of that font and renders one test caption.
- Looks at the rendered test to confirm the letters join and the line runs right to left.
- Reads every word with its timing, sources B-roll and music, locates the speaker's face and tests the punch-in framing before rendering the base cut.
Step 8 is the one that matters. Persian text rendered without shaping still produces an image, and it produces disconnected letters in the wrong order. Nothing errors. Vizard Agent looked at one frame rather than trusting the render.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 44.7 seconds, AAC audio. Vertical, forty-five seconds, the original Persian speech with correctly shaped right-to-left captions burned in, B-roll over the talk and a music bed underneath.
The run hit ten failed steps along the way, most of them in the font and shaping stretch. Vizard Agent read each error and fixed the cause rather than retrying the same command, which is why the captions came out right rather than approximately right.
Forty-five seconds is a tight cut from a full minute of talking. Vizard Agent removed the dead air first and covered the joins with B-roll, which is why the reel reads as edited rather than as a phone recording with words on it.
When does this not work well?
Right-to-left scripts are the case where "it produced something" and "it produced the correct thing" look completely identical in a log. Vizard Agent verifies by rendering a test and looking at it, and there are still real limits worth knowing before you brief the job.
- Mixed direction is genuinely hard. A Latin brand name inside a Persian sentence flips in ways that need checking by someone who reads the language.
- Numbers and punctuation sit differently. Where a comma or a digit lands is a real question in an RTL line.
- Transcription quality varies by language. A noisy car recording in a lower-resource language transcribes less accurately than clean English.
- Not every font covers the script. Vizard Agent fetches one that does, and an unusual style may simply not exist.
- Read the result yourself. If you do not read the language, ask somebody who does before publishing.
How do you fix a result that came back wrong?
Say what is wrong with the text. Vizard Agent keeps the transcript, the word timings, the installed font and the rendered test frames, so a wording fix or a restyle re-renders the captions against work it already did rather than repeating the font setup.
- "That word is transcribed wrong." Corrected in the transcript and re-burned.
- "The captions are too big." A restyle over the same cut.
- "Add an English translation underneath." A second caption track from the same timings.
How does Vizard Agent compare to doing it yourself?
By hand this is finding a font that covers the script, discovering your editor does not shape the letters, exporting a caption file that looks correct in a text editor and wrong in the video, and having no idea which of those three steps broke it.
| By hand | Vizard Agent | |
|---|---|---|
| The font | Hunt for one that covers the script | Fetched automatically |
| Letter shaping | Silently wrong | Libraries installed, then verified |
| Direction | Discover on export | Checked on a test frame |
| Caption timing | Type and nudge | From word-level timings |
Common questions
Which right-to-left languages does this work for? Persian, Arabic, Hebrew, Urdu and others. Vizard Agent fetches the font and shaping the script needs.
Does it translate the speech? Not unless you ask. By default Vizard Agent transcribes in the language spoken and captions in the same one.
Can I have two caption tracks? Yes. Ask for the original and a translation, and both come from the same word timings.
Will emoji work in the captions? Vizard Agent checks emoji support in the font before using them, so say if you want them.
Can it add B-roll over the talk? Yes. It sources clips against what is being said and cuts them over the speaker.
How do I know the text is right? Vizard Agent renders a test caption and looks at it, and if you do not read the script yourself, have someone who does check the delivery.