How to get a voiceover on its own without a video
Paste the script, describe who should be speaking, and tell Vizard Agent you want the audio on its own. Vizard Agent finds candidate voices, generates the script in several of them at once, analyses each take for accent and intonation, measures the real speaking time, and hands you the file.
What is the short version?
Sometimes the edit already exists and the only thing missing is the voice. Asking Vizard Agent for the audio on its own is a perfectly ordinary request, and it lets you audition several readings of the same script before committing any of them to a timeline.
- Go to Vizard Agent and paste the script.
- Describe the speaker — age, accent, manner, who they sound like.
- Say you want audio only, no video.
What do you need before you start?
The script and a picture of the speaker. Vizard Agent searches for voices that match a description rather than a category, so "a woman of about twenty-five with a natural Spanish accent, talking like a friend" gets you much closer than "female voice".
- The script. Word for word, as it should be read.
- The speaker. Age, accent, register.
- The relationship. Who they are talking to.
- The tone. Persuasive, calm, urgent, warm.
- The length target. If it has to fit something.
What do you type into Vizard Agent?
Describe the person, not the product. Vizard Agent uses the description to search a voice library and to choose between candidates, so details about how they talk and who they are talking to are worth more than adjectives about the brand.
Prompt
Variants worth knowing:
- Several takes compared. Different voices, same script.
- Punctuation for pauses. The script prepared so it breathes.
- A duration target. So it fits a cut you already have.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real audio-only request. The step that most people skip is the eleventh: it measured how much of each take was actual speech rather than silence, which is what tells you whether a take will fit your edit.
- Checks the voiceover options available to it.
- Searches for young female voices with the right regional accent, then searches again for more.
- Prepares the script with pauses and punctuation so the reading breathes properly.
- Generates the narration in several voices at once rather than one at a time.
- Checks the duration of each generated take.
- Analyses the intonation, accent and recording quality of every one.
- Pulls the word-level timings from the leading candidate.
- Calculates net speaking time against silence for each take.
- Generates further test takes to be sure of the tone, then analyses the remaining candidates the same way before choosing.
Step 8 is the practical one. Two takes of the same script can differ by several seconds purely in how much air is between the sentences, and knowing the real speaking time tells you which one will fit a fixed cut.
What does the result look like?
An audio file of the script read in the chosen voice, with the punctuation-driven pauses in the right places, ready to drop into whatever you are editing. There is no video in this job at all — the deliverable is the recording.
Alongside it you have the alternative takes, which is often the more useful part. Hearing the same script in three voices is how you find out that the one you assumed you wanted is the wrong one.
When does this not work well?
A generated voice is very convincing at some things and recognisably not a person doing others, and Vizard Agent cannot close that gap for you. These are the limits worth knowing honestly before you build a campaign around one.
- Regional accents are approximations. They suggest a region rather than come from one.
- Long scripts drift. Energy is harder to sustain across several minutes.
- Names and jargon get mispronounced. Spell them phonetically if they matter.
- Emotion has a ceiling. Genuine distress or laughter reads as performed.
- Cloning needs consent. Only clone a voice you are entitled to use.
How do you fix a result that came back wrong?
Say what is wrong with the reading itself. Vizard Agent keeps every take it generated, the word timings and the speaking-time measurements, so a different voice or a re-punctuated script can be produced without starting the voice search from scratch.
- "Too fast." The script is re-punctuated and regenerated with longer pauses.
- "Wrong accent." Swapped for one of the takes already generated.
- "It mispronounces the brand." Respelled phonetically and re-read.
How does Vizard Agent compare to doing it yourself?
By hand this means booking a voice artist, or recording it yourself and discovering the room sounds wrong. Either way you get one reading, and comparing alternatives means paying for or recording every one of them, which is why almost nobody does.
| By hand | Vizard Agent | |
|---|---|---|
| Casting | Listen to demo reels | Several takes of your actual script |
| Comparing | One version, usually | Every candidate analysed |
| Fitting a cut | Time it and re-record | Net speaking time measured per take |
| Turnaround | Days | The same session |
Common questions
Can I really get just the audio? Yes. Vizard Agent delivers the recording as a file with no video attached.
How many voices will it try? Several at once. Vizard Agent generated multiple takes here and analysed every one of them.
Can I hear the alternatives? Yes. Vizard Agent keeps all the takes it generated so you can compare the readings yourself.
How specific should the description be? Very. Age, accent, register and who they are talking to all change what Vizard Agent picks.
Will it fit my existing edit? Vizard Agent measures the net speaking time, so you can pick the take that actually fits.
Why measure speech against silence? Because two takes of identical wording can run several seconds apart purely on pause length. If the voice has to land in a fixed cut, the total duration tells you less than how much of it is actually speech.
Can it read in any language? Yes. Vizard Agent finds a voice suited to it — say the region as well as the language.
Can I use my own cloned voice? Yes. Vizard Agent uses a cloned voice provided it is yours or you have permission.
What if it mispronounces something? Respell it phonetically in the script and Vizard Agent reads it correctly.
Can I get the word timings with it? Yes, and they are worth asking for. Vizard Agent already extracts word-level timings to measure the takes, so having them alongside the audio means your captions and cuts can be built against the actual reading.
Should I write the script differently for a generated voice? A little. Vizard Agent prepares the punctuation so the reading breathes, but short sentences and plain phrasing carry better than long written-for-the-page constructions do.
Can Vizard Agent build the video afterwards too? Yes. Asking for audio only keeps the two steps separate, and once you have chosen a take it can go straight on to cutting picture against those same word timings.
Does Vizard Agent check the recording? Yes. It analyses the accent, intonation and quality of each take before choosing one.