How to narrate a video in your own voice
Upload a clean sample of your voice to Vizard Agent and paste the script you want read. Vizard Agent checks the sample, crops the cleanest stretch out of it, generates the narration section by section in your voice, joins the parts and checks the waveform before handing the file over.
What is the short version?
Recording narration is the step that stops videos getting made: you need a quiet room, an hour, and the patience not to fluff a line for two minutes at a time. A voice sample removes all three, and the quality of what comes back is set almost entirely by the quality of what you sent.
- Go to Vizard Agent and upload a sample of yourself speaking.
- Paste the script exactly as you want it read.
- Say the delivery you want, then listen to the whole thing once.
What do you need before you start?
A clean sample and a final script. Both requirements are stricter than they sound: Vizard Agent learns whatever is in the sample, including the room, and it reads exactly the characters you paste, including the ones you meant to say differently.
- 10 to 90 seconds of clean speech. One speaker, no music, no second voice. This is a hard requirement, not a preference.
- The quietest recording you have. Room tone, hiss and echo are learned along with your voice.
- The script as final text. Numbers, names and abbreviations are read as written, so write "twenty twenty six" if that is what you want heard.
- A note on delivery. Warm, brisk, serious — and where the pauses fall, if they matter.
What do you type into Vizard Agent?
Paste the script in full rather than describing it. Vizard Agent generates from the exact characters you send, so anything left to interpretation is a place where the narration can differ from what you had in your head — and unlike a picture, you will not notice until you listen the whole way through.
Prompt
Variants worth knowing:
- Audio only. Say "audio file only" and Vizard Agent returns the narration for you to use elsewhere.
- Straight onto the video. Upload the footage too and ask Vizard Agent to lay the narration over it.
- Another language in your voice. Say which. That is a translation job and a narration job at once, so tell Vizard Agent both halves.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real narration job. Two of the steps are about your sample rather than your script, and those are the ones that decide how the result sounds — everything after them is execution, and execution is not where cloned voices usually go wrong.
- Reads the voice generation tool's own documentation before choosing any settings.
- Confirms whose voice is in the sample and what is being said in it.
- Converts and uploads the voice sample.
- Analyses the sample's quality.
- Crops a clean stretch out of it rather than using the whole file, because the clean part is the part worth learning.
- Generates the narration section by section, not in a single pass.
- Joins the sections and checks the waveform before uploading the result.
Step 5 is the one that matters most. A sample with a cough in it teaches the cough, and Vizard Agent cutting around it is the difference between a voice that sounds like you and one that sounds like your worst take.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 52.48 seconds, AAC audio. That run put the narration onto video; ask Vizard Agent for audio only and you get the audio file instead.
Voiceover is the most expensive category to run, at roughly 200 credits per finished minute, so it is worth getting the script final before the first generation rather than after the third.
When does this not work well?
Two things break this: a sample that is not clean, and a script that assumes the reader knows what you meant. Vizard Agent can tell you the first has gone wrong; the second only shows up when you listen.
- A sample with music under it will not work. Nor will one with two people talking. The requirement is a clean solo stretch, and a sample that fails it produces a voice that is nearly you and wrong in a way that is hard to name.
- It reads what you wrote. Homographs are guesswork: "read", "lead", "live". If a word matters, spell the pronunciation you want.
- Strong regional accent and dialect are harder. More sample helps. Check the result rather than trusting it.
- One language per run. A bilingual script is two runs, not one, and Vizard Agent will say so rather than half-doing it.
- A cloned voice cannot ad-lib. It reads the script and nothing else, so the small unscripted noises that make speech sound live — a breath before a hard word, a laugh in the middle of a sentence — are not there unless you write them in.
- Consent is yours to hold. Whether Vizard Agent can clone a voice and whether you may are different questions. Use your own, or one you have written permission for.
How do you fix a result that came back wrong?
Point at the line. Vizard Agent generates the narration in sections and keeps them, so a correction regenerates the section you named rather than the whole file — which matters here more than elsewhere, because this is the priciest thing to redo.
- "The third paragraph is too fast." That section is regenerated at the pace you give.
- "It said the name wrong." Spell it phonetically and only that line changes.
- "Warmer overall." A delivery change, applied across the sections.
How does Vizard Agent compare to doing it yourself?
Recording it yourself is free and sounds exactly like you, which is a real advantage worth keeping in mind. What it costs is a quiet hour, a microphone, and doing the whole take again when you change one sentence in the script — which, on a video that goes through three drafts, is three recording sessions.
| Recording it yourself | Vizard Agent | |
|---|---|---|
| Setup | Quiet room, microphone | A 30-second sample, once |
| Changing a line | Re-record and re-cut | Regenerates that section |
| Consistency across videos | Depends on the day you had | Same voice every time |
| Another language | Not an option for most people | Say which language |
Common questions
How long a sample do I need? Between 10 and 90 seconds of clean single-speaker audio. Longer is not better if it is noisier.
Can I get the audio without a video? Yes. Ask Vizard Agent for "audio file only".
Will it match my pacing? It follows the sample's character and whatever delivery you state. If the pace is wrong, name the section and Vizard Agent regenerates that part.
Does the sample get stored? It stays in the project you uploaded it to, the same as any other file you send. If you want it out, say so and Vizard Agent will delete it from the workspace.
Can it lip-sync my face to it afterwards? Yes, and it is a common next step. Upload a photo or clip of yourself in the same project so the narration is already there to sync against.