Vizard Agent

How to get the on-screen text out of a video as a document

Last updated 2026-08-25 · 7 min read

Upload the video and tell Vizard Agent you want the text as a document. Vizard Agent extracts frames across the whole recording, tiles them so the pages can be read, crops and enlarges each block to check it word by word, and writes the result out as a file you can edit.

What is the short version?

Sometimes what you need out of a video is not a video. A recorded lecture with slides, a screen recording of a document, a presentation someone sent as an MP4 — the value is the text, and retyping it by hand is the reason it never gets used.

  1. Go to Vizard Agent and upload the video.
  2. Say you want the on-screen text as a document, not a video.
  3. Say which format — a Word file, or plain text.

What do you need before you start?

The video and a decision about what counts as text. Vizard Agent can pull both what is written on screen and what is spoken aloud, and saying which you want avoids getting a document with two overlapping versions of the same material in it.

What do you type into Vizard Agent?

Say the output is a document. That one word changes what Vizard Agent builds — a file rather than a render — and it will install whatever it needs to write that format rather than handing you a transcript pasted into a message.

Prompt

Convert the text in this video into a [Word] document. Keep the on-screen text as written, in the order it appears. [Include / exclude] what is spoken aloud.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real video-to-document job. The step that decides accuracy comes late: rather than trusting one reading of the frames, it cropped each text block and enlarged it to check the words individually.

  1. Checks the video's details and extracts text from both the speech and the frames.
  2. Tiles the frames into a composite so the content can be seen at a glance.
  3. Looks at the tiled frames to work out where the text actually lives.
  4. Checks which document libraries are available and installs one that writes Word files.
  5. Extracts detailed frames of the document sections rather than sampling evenly.
  6. Joins the document frames together so the whole content can be read in order.
  7. Reads the text through in sections, from the opening to the closing frames.
  8. Crops the sharpest text regions for a closer look at the wording.
  9. Builds enlarged groups of each block to check it word by word, then writes the document out.

Step 9 is the difference between a document you can use and one you have to proofread against the video. Text read at frame resolution is text with plausible-looking errors in it, and enlarging each block is how those get caught.

What does the result look like?

A document file containing the video's on-screen text in the order it appeared, structured as it was laid out rather than as one unbroken block, ready to edit, search and paste from. The video itself is untouched and unneeded from that point on.

There is no rendered video in this job at all. That is worth saying plainly, because it is the one request in this collection where the deliverable is a file you open in a word processor.

When does this not work well?

Reading text off video frames has a hard floor set by the recording itself, and no amount of care from Vizard Agent raises it. What it does instead is report the sections it could not read, rather than inventing plausible words to fill the gap.

How do you fix a result that came back wrong?

Point at the section that came out wrong. Vizard Agent keeps the extracted frames, the tiled composites and the enlarged crops, so a passage can be re-read at higher magnification without processing the whole video a second time.

How does Vizard Agent compare to doing it yourself?

By hand this means pausing the video repeatedly, screenshotting, and typing what you see, or running frames through a text recogniser and then proofreading every line of it. Both are slow, and the second one produces errors that read as correct.

By hand Vizard Agent
Capturing Pause and screenshot Frames extracted across the whole video
Reading Type it out Read from tiled composites in order
Accuracy Proofread against the video Each block enlarged and checked
Output Paste into a document Written directly as a file

Common questions

Can Vizard Agent get spoken words as well? Yes. It extracts both, and keeps them as separate sections when you ask.

What formats can it write? Word files and plain text. Vizard Agent installs what it needs to write the format you name.

Does the video quality matter? Considerably. Vizard Agent reads what is legible, and compression blurs small text first.

Will the structure be preserved? Yes. Vizard Agent keeps headings, lists and paragraph breaks as they appeared on screen.

What about text in another language? Say which. Vizard Agent reads it, and knowing the language makes the reading more reliable.

Why enlarge each block instead of reading once? Because text read at frame resolution produces errors that look entirely plausible — a wrong digit, a swapped word — and those survive proofreading. Enlarging each block is what catches them before they reach the document.

How long can the video be? Any length. Longer recordings simply mean more frames for Vizard Agent to read through.

Can it handle slides and a speaker together? Yes. Vizard Agent finds where the text lives in the frame rather than assuming it fills it.

Am I allowed to extract text from any video? Extraction is not permission. Someone else's material remains theirs regardless of format.

Can I get the slides as images too? Yes. Vizard Agent has already extracted and tiled the frames, so saving the clean full-size slide images alongside the document costs almost nothing extra.

Does this work on a screen recording? Yes, and it is close to the ideal case. Screen recordings are sharp and the text is rendered rather than filmed, so Vizard Agent reads them far more reliably than a camera pointed at a projector.

Does Vizard Agent tell me what it could not read? Yes. It reports the sections it could not resolve rather than guessing at them.