How to get a readable transcript instead of a verbatim one
Ask Vizard Agent for a clean transcript rather than a verbatim one, and say what to leave out. Verbatim is the default because it is the safe answer for subtitles and legal records — but as a document to read, it is almost unusable, and cleaning it is a separate instruction.
What is the short version?
You want the text of a recorded lesson as something a person can actually sit and read. Word-for-word gives you every hesitation, every false start and three minutes of housekeeping at the top of it. Vizard Agent edits the transcript for sense instead of transcribing it literally.
- Ask Vizard Agent for a clean transcript, not verbatim.
- Name what should not appear at all.
- Say how it should be broken up and headed.
What do you need before you start?
The recording, and a short list of exclusions. In the session behind this article the exclusions were specific — the speaker's own introduction and the administrative information about the course — and naming them is what stopped the document opening on material nobody reading it needed.
Decide the structure too. Paragraphs by topic, a heading with the lesson number, nothing else: that is a complete specification, and it takes one sentence. Without it you get a wall of text that is clean and still hard to use.
What do you type into Vizard Agent?
Tell Vizard Agent three things: that you want it clean, what to leave out entirely, and what shape the document should have. All three fit into a couple of lines, and each one of them changes the output noticeably.
Make a clean transcript with the filler removed. Leave out the speaker's introduction and the course admin. Keep the meaning exact, break it into paragraphs by topic, and put the lesson number in the heading.
"Keep the meaning exact" is the important counterweight. Cleaning a transcript means removing words, and the line between removing filler and paraphrasing is one you should draw explicitly.
What does Vizard Agent actually do?
It transcribes and then edits, as two passes. The raw text is prepared without timestamps, because timestamps are what make a transcript unreadable as prose, and then it is read in sections and edited for sense rather than run through a filler-word filter.
In that session the work ran:
- Pulled frames from across the recording to see what the lesson covered.
- Read the opening phrases and checked the title slides to identify the lesson and its number.
- Prepared the transcript without timestamps, in a form convenient for editorial work.
- Read it through in four parts, editing each for meaning rather than for word count.
- Assembled the clean version into a text file and delivered it as a download.
Reading it through in parts is not a formality. A filler-word filter produces something that still reads unmistakably like speech, because the problem is not the individual words so much as the shape of the sentences around them; editing section by section is what turns a spoken ramble into something that reads as a written sentence.
What does the result look like?
A document. Paragraphs that each cover one thing, a heading that tells you which lesson it is, no timestamps, no speaker labels unless the recording needs them, and prose that reads as though it were written rather than said.
The heading comes from the recording rather than from the file. Reading the title slide is how the lesson number gets into the document correctly — filenames drift, and a document headed with the wrong lesson number is worse than one with no heading.
When does this not work well?
When you need the exact words. A quote, a legal record, a subtitle file and a compliance review all want verbatim, and a cleaned transcript is not a substitute. Ask for both if you need both; they come from the same pass.
It is also a poor fit for conversation. Two people interrupting each other cleans into something that reads oddly, because the mess is the shape of the exchange. A summary usually serves better there than a cleaned transcript.
And it cannot rescue unclear audio. Where the recognition is uncertain, cleaning the text makes that uncertainty less visible rather than less real, which is worse than leaving it rough — so say if the audio is poor, and Vizard Agent will flag the doubtful passages instead of smoothing them into confident prose.
How do you fix a result that came back wrong?
If the meaning shifted, quote the sentence. Cleaning is the one transcript operation that can change what someone said, so a specific example is worth more than a general note, and Vizard Agent can go back to the original words for that passage.
If too much came out, say what you missed. "Keep the examples" and "keep the asides" are different instructions, and either can be applied without redoing the whole document.
If the structure is wrong, describe the shape rather than editing it. Paragraph breaks and headings are cheap to regenerate and tedious to fix by hand.
How does Vizard Agent compare to doing it yourself?
Any transcription tool gives you verbatim text in a minute. Turning that into something readable is the part that takes the afternoon, and it is genuinely editorial work — deciding what was filler and what was emphasis is a judgement, not a find-and-replace.
What Vizard Agent does is that second pass, in sections, against the recording rather than against the text alone. And on a course, it does it the same way for every lesson — which is what makes a set of them usable as a set.
Common questions
Can I have both versions? Yes. Vizard Agent produces the verbatim text on the way to the clean one.
Will it change what was said? It should not. Say "keep the meaning exact" and Vizard Agent removes rather than rewrites.
What counts as filler? Hesitations, restarts and throat-clearing. Tell Vizard Agent if you want more or less removed.
Can it keep the timestamps? Yes if you ask Vizard Agent for them, though they make it much harder to read as prose.
What about speaker names? Vizard Agent includes them when there is more than one speaker. Say if you want them out.
Can it do a whole course? Yes. Vizard Agent applies the same treatment to each lesson.
Where does the heading come from? The title slide in the recording rather than the filename.
What format do I get? A text file by default. Ask Vizard Agent for another if you need one.
Can it summarise as well? Yes. Vizard Agent produces it as a separate output alongside the transcript.
What if the audio is poor? Say so and Vizard Agent flags the uncertain passages rather than smoothing them.
Can it translate the clean version? Yes, and translating the cleaned text gives a better result than the raw one.
Does it keep the original order? Yes, unless you ask Vizard Agent to restructure it by topic.
How long does it take? Minutes. For Vizard Agent the editing pass is the slower half, not the transcription.
Can I give it my own style rules? Yes. Vizard Agent follows them, including conventions like spelling out numbers or keeping technical terms untranslated.
Can it split one recording into several documents? Yes. Ask Vizard Agent for one file per topic and it breaks them at the topic boundaries.