Vizard Agent

How to have the text on screen read aloud in your video

Last updated 2026-08-31 · 8 min read

Upload the video and ask for the on-screen text to be narrated. Vizard Agent reads the captions off the frames themselves, detects the exact moment each one appears and disappears, records a line for each, and delays every one to land with its caption.

What is the short version?

A silent video full of captions works on a feed and nowhere else. Adding a voice that reads them turns it into something that works on YouTube, in an email, or for anyone who cannot easily read text on a small screen.

  1. Go to Vizard Agent and upload the video.
  2. Say that the on-screen captions should be read aloud.
  3. Say what the voice should sound like.

What do you need before you start?

Only the video. The awkward part is that a video with burnt-in captions has no subtitle file and no transcript, so there is nothing to hand over except the picture — and Vizard Agent has to read the words out of the frames to work out what to say.

What do you type into Vizard Agent?

Say what to read and how it should sound. Vizard Agent works out the timing itself by detecting when each caption appears, so the instruction that matters is about the voice and about what happens to the audio already in the video.

Prompt

Add a voiceover that reads the on-screen captions aloud through the whole video. The voice should be [clear and professional], and each line should be synchronised to the caption it belongs to — [narrating the before and after]. Keep the original audio underneath at a lower level.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real thirty-six-second video with burnt-in captions and no subtitle file. Everything up to step twenty-one is the detective work of finding out what the captions say and precisely when each one is on screen.

  1. Probes the video for resolution, duration, frame rate and audio.
  2. Extracts frames and tiles them into a grid to look for on-screen text.
  3. Reads the grid and identifies the captions visually.
  4. Checks the workspace for any existing script or subtitle file — there is none.
  5. Runs shot transcription and speech recognition to see whether any text already exists.
  6. Writes a script to detect the exact timestamps at which the captions change.
  7. Crops the text area every half second with timestamps, to see each caption's window.
  8. Locates the precise bounding box of the text in a sample frame.
  9. Narrows the crop to the band where the text actually sits and regenerates the grid.
  10. Verifies every caption's wording and its appearance and disappearance times.
  11. Searches the voice library for a suitable professional narrator.
  12. Generates the eleven narration segments in parallel and measures each one's duration.
  13. Measures the original audio's loudness and checks the voice's channel layout.
  14. Delays each segment to its caption's start and mixes it over the original audio.
  15. Verifies the render's sync, narration clarity and background balance.

Steps six to ten are the whole job. There is no subtitle file to read, so Vizard Agent has to establish both the text and its timing from the pixels — and getting the crop band right is what turns a smear of half-captions into a clean list of lines with times against them.

What does the result look like?

The video this page is written from measured 1080x1920 at 30fps and ran 36.43 seconds with existing audio. Eleven captions came off the screen with their exact windows, and eleven narration segments were generated and delayed into place.

What comes back sounds deliberate rather than dubbed. Each line starts as its caption appears, the original audio continues underneath at a lower level, and nothing overlaps because the segment lengths were measured before placement.

When does this not work well?

Reading text out of pixels is harder than reading it out of a file, and some videos make it much harder still. Vizard Agent tells you which captions it could not resolve rather than guessing at a word and narrating something you never wrote.

How do you fix a result that came back wrong?

Give the corrected line and its number. Vizard Agent keeps the caption list, the detected timings and the recorded segments, so a wrong word or a late entry is a re-record of one line rather than the whole track.

How does Vizard Agent compare to doing it yourself?

By hand this means pausing the video eleven times to type out the captions, recording or generating each line, and then dragging every one onto a timeline until it lines up. Vizard Agent detects the transitions programmatically and delays each segment to a measured start.

By hand Vizard Agent
The wording Typed out by pausing Read from the frames
The timing Dragged into place Detected to the timestamp
Segment lengths Discovered when they overrun Measured before placement
The mix Levelled by ear Original loudness measured first

Common questions

Do I need a subtitle file? No, and that is the point. Vizard Agent reads the captions off the picture itself.

What if the video has audio already? It stays. Vizard Agent mixes the narration over it at a measured level.

Can I choose the voice? Yes. Vizard Agent searches the voice library against your description.

Will the lines overlap? Not if the timings allow. Vizard Agent measures each segment before placing it.

What if a caption is too long to read? It tells you. Vizard Agent flags lines that will not fit in their window.

Does it work with animated captions? Less reliably. Vizard Agent handles static lines far better than word-by-word effects.

Can it narrate in another language? Yes. Vizard Agent reads the captions and records the narration in the language you name.

Is this useful for accessibility? Very. Vizard Agent turns a caption-only video into something that works without reading.

Can it also produce a text file? Yes. Vizard Agent can output the caption list with its timings alongside the video.

How long does it take? Minutes for a short video. Vizard Agent generates the segments in parallel.

Why not just retype the captions myself? You can, and for one video that is fine. The value is in the timings — Vizard Agent detects each caption's exact frame, which is the part that takes a person the longest to get right.

What if the text moves around the frame? Say where it appears. Vizard Agent narrows its crop to the right region rather than scanning the whole frame.

Does it check the result? Yes. Vizard Agent verifies the sync, the clarity of the narration and the balance against the background audio.