How to have the text on screen read aloud in your video
Upload the video and ask for the on-screen text to be narrated. Vizard Agent reads the captions off the frames themselves, detects the exact moment each one appears and disappears, records a line for each, and delays every one to land with its caption.
What is the short version?
A silent video full of captions works on a feed and nowhere else. Adding a voice that reads them turns it into something that works on YouTube, in an email, or for anyone who cannot easily read text on a small screen.
- Go to Vizard Agent and upload the video.
- Say that the on-screen captions should be read aloud.
- Say what the voice should sound like.
What do you need before you start?
Only the video. The awkward part is that a video with burnt-in captions has no subtitle file and no transcript, so there is nothing to hand over except the picture — and Vizard Agent has to read the words out of the frames to work out what to say.
- The video. With its text already burnt in.
- The voice. Clear, professional, or a described character.
- The mix. Whether the original audio stays underneath.
- The pace. Whether lines can run over each other.
- The intent. Narration, or accessibility, or both.
What do you type into Vizard Agent?
Say what to read and how it should sound. Vizard Agent works out the timing itself by detecting when each caption appears, so the instruction that matters is about the voice and about what happens to the audio already in the video.
Prompt
Variants worth knowing:
- Say "synchronised to the captions." Not just "add narration".
- Say what happens underneath. Original audio kept or dropped.
- Ask for a matching voice. If there is already one in the file.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real thirty-six-second video with burnt-in captions and no subtitle file. Everything up to step twenty-one is the detective work of finding out what the captions say and precisely when each one is on screen.
- Probes the video for resolution, duration, frame rate and audio.
- Extracts frames and tiles them into a grid to look for on-screen text.
- Reads the grid and identifies the captions visually.
- Checks the workspace for any existing script or subtitle file — there is none.
- Runs shot transcription and speech recognition to see whether any text already exists.
- Writes a script to detect the exact timestamps at which the captions change.
- Crops the text area every half second with timestamps, to see each caption's window.
- Locates the precise bounding box of the text in a sample frame.
- Narrows the crop to the band where the text actually sits and regenerates the grid.
- Verifies every caption's wording and its appearance and disappearance times.
- Searches the voice library for a suitable professional narrator.
- Generates the eleven narration segments in parallel and measures each one's duration.
- Measures the original audio's loudness and checks the voice's channel layout.
- Delays each segment to its caption's start and mixes it over the original audio.
- Verifies the render's sync, narration clarity and background balance.
Steps six to ten are the whole job. There is no subtitle file to read, so Vizard Agent has to establish both the text and its timing from the pixels — and getting the crop band right is what turns a smear of half-captions into a clean list of lines with times against them.
What does the result look like?
The video this page is written from measured 1080x1920 at 30fps and ran 36.43 seconds with existing audio. Eleven captions came off the screen with their exact windows, and eleven narration segments were generated and delayed into place.
What comes back sounds deliberate rather than dubbed. Each line starts as its caption appears, the original audio continues underneath at a lower level, and nothing overlaps because the segment lengths were measured before placement.
When does this not work well?
Reading text out of pixels is harder than reading it out of a file, and some videos make it much harder still. Vizard Agent tells you which captions it could not resolve rather than guessing at a word and narrating something you never wrote.
- Text over busy footage. Hard to isolate cleanly.
- Decorative or animated type. Word-by-word effects confuse it.
- Very fast captions. A line may not fit as speech.
- Text in several places. Say where the captions sit.
- Stylised fonts. Legibility for a reader is not legibility for a crop.
How do you fix a result that came back wrong?
Give the corrected line and its number. Vizard Agent keeps the caption list, the detected timings and the recorded segments, so a wrong word or a late entry is a re-record of one line rather than the whole track.
- "Line six says the wrong word." Re-read and re-placed.
- "The voice starts too early." Delayed to the caption's exact frame.
- "The music drowns it." Original audio lowered under the narration.
How does Vizard Agent compare to doing it yourself?
By hand this means pausing the video eleven times to type out the captions, recording or generating each line, and then dragging every one onto a timeline until it lines up. Vizard Agent detects the transitions programmatically and delays each segment to a measured start.
| By hand | Vizard Agent | |
|---|---|---|
| The wording | Typed out by pausing | Read from the frames |
| The timing | Dragged into place | Detected to the timestamp |
| Segment lengths | Discovered when they overrun | Measured before placement |
| The mix | Levelled by ear | Original loudness measured first |
Common questions
Do I need a subtitle file? No, and that is the point. Vizard Agent reads the captions off the picture itself.
What if the video has audio already? It stays. Vizard Agent mixes the narration over it at a measured level.
Can I choose the voice? Yes. Vizard Agent searches the voice library against your description.
Will the lines overlap? Not if the timings allow. Vizard Agent measures each segment before placing it.
What if a caption is too long to read? It tells you. Vizard Agent flags lines that will not fit in their window.
Does it work with animated captions? Less reliably. Vizard Agent handles static lines far better than word-by-word effects.
Can it narrate in another language? Yes. Vizard Agent reads the captions and records the narration in the language you name.
Is this useful for accessibility? Very. Vizard Agent turns a caption-only video into something that works without reading.
Can it also produce a text file? Yes. Vizard Agent can output the caption list with its timings alongside the video.
How long does it take? Minutes for a short video. Vizard Agent generates the segments in parallel.
Why not just retype the captions myself? You can, and for one video that is fine. The value is in the timings — Vizard Agent detects each caption's exact frame, which is the part that takes a person the longest to get right.
What if the text moves around the frame? Say where it appears. Vizard Agent narrows its crop to the right region rather than scanning the whole frame.
Does it check the result? Yes. Vizard Agent verifies the sync, the clarity of the narration and the balance against the background audio.