Vizard Agent

How to narrate a video in your own voice

Last updated 2026-08-16 · 5 min read

Upload a clean sample of your voice to Vizard Agent and paste the script you want read. Vizard Agent checks the sample, crops the cleanest stretch out of it, generates the narration section by section in your voice, joins the parts and checks the waveform before handing the file over.

What is the short version?

Recording narration is the step that stops videos getting made: you need a quiet room, an hour, and the patience not to fluff a line for two minutes at a time. A voice sample removes all three, and the quality of what comes back is set almost entirely by the quality of what you sent.

  1. Go to Vizard Agent and upload a sample of yourself speaking.
  2. Paste the script exactly as you want it read.
  3. Say the delivery you want, then listen to the whole thing once.

What do you need before you start?

A clean sample and a final script. Both requirements are stricter than they sound: Vizard Agent learns whatever is in the sample, including the room, and it reads exactly the characters you paste, including the ones you meant to say differently.

What do you type into Vizard Agent?

Paste the script in full rather than describing it. Vizard Agent generates from the exact characters you send, so anything left to interpretation is a place where the narration can differ from what you had in your head — and unlike a picture, you will not notice until you listen the whole way through.

Prompt

Use my own voice from the attached sample to narrate the following script. [paste script]. Delivery: calm and warm, with a pause before the last line.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real narration job. Two of the steps are about your sample rather than your script, and those are the ones that decide how the result sounds — everything after them is execution, and execution is not where cloned voices usually go wrong.

  1. Reads the voice generation tool's own documentation before choosing any settings.
  2. Confirms whose voice is in the sample and what is being said in it.
  3. Converts and uploads the voice sample.
  4. Analyses the sample's quality.
  5. Crops a clean stretch out of it rather than using the whole file, because the clean part is the part worth learning.
  6. Generates the narration section by section, not in a single pass.
  7. Joins the sections and checks the waveform before uploading the result.

Step 5 is the one that matters most. A sample with a cough in it teaches the cough, and Vizard Agent cutting around it is the difference between a voice that sounds like you and one that sounds like your worst take.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 52.48 seconds, AAC audio. That run put the narration onto video; ask Vizard Agent for audio only and you get the audio file instead.

Voiceover is the most expensive category to run, at roughly 200 credits per finished minute, so it is worth getting the script final before the first generation rather than after the third.

When does this not work well?

Two things break this: a sample that is not clean, and a script that assumes the reader knows what you meant. Vizard Agent can tell you the first has gone wrong; the second only shows up when you listen.

How do you fix a result that came back wrong?

Point at the line. Vizard Agent generates the narration in sections and keeps them, so a correction regenerates the section you named rather than the whole file — which matters here more than elsewhere, because this is the priciest thing to redo.

How does Vizard Agent compare to doing it yourself?

Recording it yourself is free and sounds exactly like you, which is a real advantage worth keeping in mind. What it costs is a quiet hour, a microphone, and doing the whole take again when you change one sentence in the script — which, on a video that goes through three drafts, is three recording sessions.

Recording it yourself Vizard Agent
Setup Quiet room, microphone A 30-second sample, once
Changing a line Re-record and re-cut Regenerates that section
Consistency across videos Depends on the day you had Same voice every time
Another language Not an option for most people Say which language

Common questions

How long a sample do I need? Between 10 and 90 seconds of clean single-speaker audio. Longer is not better if it is noisier.

Can I get the audio without a video? Yes. Ask Vizard Agent for "audio file only".

Will it match my pacing? It follows the sample's character and whatever delivery you state. If the pace is wrong, name the section and Vizard Agent regenerates that part.

Does the sample get stored? It stays in the project you uploaded it to, the same as any other file you send. If you want it out, say so and Vizard Agent will delete it from the workspace.

Can it lip-sync my face to it afterwards? Yes, and it is a common next step. Upload a photo or clip of yourself in the same project so the narration is already there to sync against.