How to caption a language that does not put spaces between words
Tell Vizard Agent the captions have come back as fragments rather than words. It opens the subtitle file itself to see whether the text was written that way or the renderer is splitting it, rebuilds the captions at phrase level so each glyph sequence stays intact, and checks whether a word tokeniser for that script is even available before trusting any break it makes.
What is the short version?
Captioning tools break text at spaces. Thai, Lao, Khmer, Japanese and Chinese do not use them, so a tool that expects spaces either produces one caption per character or breaks a word in the middle of a shaped cluster. Both are unreadable and both look like a font problem.
- Go to Vizard Agent with the clips.
- Say the captions are coming out as fragments.
- Ask for whole phrases rather than word-by-word.
What do you need before you start?
The clips and the language. It is also worth saying what you can see — "one character at a time" and "the letters are in the wrong order" are different faults with different causes, and describing which one you have saves a round of diagnosis.
- The clips needing captions.
- The language, by name.
- What you can see. Fragments, wrong order, boxes.
- Whether the captions are essential. Sometimes they are not.
- The headline or title, if there is one.
What do you type into Vizard Agent?
Describe what is on screen rather than asking for a style change. A caption track split at the wrong level will survive every style you try, because the styling is applied to whatever chunks the text was already broken into.
Prompt
Variants worth knowing:
- "As isolated characters." Names the level of the fault.
- "Whole spoken phrases." The unit you want.
- "Check the file itself." Where the truth is.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real batch of five vertical clips captioned in Thai. It takes several attempts, and the most useful step is the one where it stops looking at the picture.
- Inspects the word-timed transcript so each clip starts on a clean beat.
- Checks the caption tool's options for this script.
- Finds an installed typeface with glyph support for the language.
- Generates the caption files and checks them when the first batch comes back empty.
- Fixes the input field and regenerates one file to verify the pipeline.
- Renders a short test before committing the whole batch.
- Inspects that test for framing, headline placement and readability.
- Improves the grouping so captions read as phrases, not fragments.
- Prepares cleaner one-line chunks and tests them.
- Inspects the shaping and the line breaks at full resolution.
- Opens one subtitle file directly to see whether the renderer is splitting it.
- Builds phrase-level files that keep each glyph sequence complete.
- Checks whether a word tokeniser for the script is installed at all.
- Renders the batch without the faulty caption layer when it is not.
- Resolves each clip's start and end so they land on complete sentences.
Step eleven is the step that ends the guessing. A caption can look fragmented because the file is fragmented or because the renderer is breaking it on the way to the screen, and those need opposite fixes — reading the file tells you which one you have in a few seconds.
Step thirteen and fourteen are the honest ending. If nothing on the system can tell where the words are in that script, then any break is a guess, and delivering the clips with a strong headline and no caption layer is better than delivering unreadable text.
What does the result look like?
Captions that read as phrases a person would say, with the characters correctly shaped and joined, breaking only where the language actually has a boundary — or, where that cannot be guaranteed, clips that carry the narration and the headline without a caption track pretending to work.
Either way, nothing on screen is unreadable.
When does this not work well?
Word boundaries in these scripts are genuinely hard to determine. Without a proper tokeniser for the language installed, every break Vizard Agent makes is an estimate, and an estimate that lands inside a word is exactly the fault you came here to fix.
- No tokeniser available. Every break is a guess.
- Mixed scripts in one line. Two sets of rules at once.
- Very long phrases. They will not fit one line.
- Typefaces without full coverage. Missing glyph boxes appear.
- Karaoke-style highlighting. It works at word level, which is the problem.
How do you fix a result that came back wrong?
Say what the text is actually doing on screen rather than how it looks. Vizard Agent keeps every subtitle file it wrote, so a fix changes how the phrases were grouped rather than re-transcribing all the audio again.
- "Still fragmented." Rebuilt at phrase level and the file checked.
- "The characters look wrong." A typeface with full coverage.
- "The line is too long." Split at a real phrase boundary.
- "Drop the captions." Delivered with the headline instead.
How does Vizard Agent compare to doing it yourself?
By hand this looks like a font problem, so the font gets changed, and then the size, and then the style — none of which touch the level the text was broken at. Opening the subtitle file is the obvious diagnostic and almost nobody does it.
| By hand | Vizard Agent | |
|---|---|---|
| First assumption | A font problem | Read the subtitle file |
| The unit | Whatever the tool produced | Rebuilt at phrase level |
| Glyph shaping | Hoped for | Checked at full resolution |
| Word boundaries | Estimated | A tokeniser, or no captions |
| When it cannot be done | Ship it anyway | Ship without the caption layer |
Common questions
Why does it come out as single characters? Because the tool breaks at spaces and your language does not use them; Vizard Agent groups by phrase instead.
Is it a font problem? Usually not. Vizard Agent reads the file before blaming the font.
Can it caption Thai or Japanese properly? Yes at phrase level, and Vizard Agent checks the shaping.
What is phrase level? A whole spoken phrase per caption from Vizard Agent, rather than a word or a character.
Can I still have word highlighting? Only where word boundaries are reliable; Vizard Agent will say.
What if no tokeniser exists? Vizard Agent delivers without the caption layer rather than guess.
Will the headline still be there? Yes. Vizard Agent keeps the headline and narration regardless.
Can it check a single file for me? Yes, and Vizard Agent finds it the fastest diagnosis available.
Does this affect the timing? No. Vizard Agent changes only how the text is grouped.
What about mixed English and Thai? Harder. Vizard Agent applies each script's rules where it can.
Will the glyphs join correctly? Vizard Agent verifies the shaping at full resolution.
Can I supply my own subtitle file? Yes, and that removes the grouping problem entirely.
Does it work across a batch? Yes. Vizard Agent applies the same treatment to every clip.
Why read the file instead of looking at the video? Because a fragmented file and a renderer that fragments look identical on screen and need opposite fixes.