How to make a reusable presenter clone of yourself
Upload footage of yourself talking and give Vizard Agent the script you want said. Vizard Agent transcribes what you recorded, picks the cleanest stretch as a voice sample, clones your voice for the new words, re-syncs your footage to them, and checks the mouth close up.
What is the short version?
Recording yourself once and reusing the result is the appeal here, and it works because the two halves come apart cleanly. Your voice becomes a model that can say new words, and the footage you already have gets re-synced to whatever that model says.
- Go to Vizard Agent and upload footage of yourself talking to camera.
- Add a photo or two if you have them.
- Give the script you want the clone to say.
What do you need before you start?
Footage of yourself and something for the clone to say. Vizard Agent picks the usable stretch out of what you upload, so a perfect take is unnecessary, but the recording does need a section where you are speaking clearly and nothing else is.
- The footage. You, talking to camera.
- A clean stretch. Even thirty seconds helps.
- The new script. What the clone should say.
- Photos. Optional, for stills and thumbnails.
- The use. Which videos this will front.
What do you type into Vizard Agent?
Say it is you and give the script. Vizard Agent treats a request to clone the person in the uploaded footage differently from a request to generate a presenter, and the distinction matters because one of them is your own likeness and the other is not.
Prompt
Variants worth knowing:
- A voice model kept. So later scripts do not need new footage.
- Another language. Your voice, different words.
- A short reusable clip. Rather than one long video.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real presenter clone. The step that separates a usable clone from an uncanny one is the eighteenth: it zoomed into the mouth rather than judging the lip-sync from the full frame.
- Checks the project history and which avatar tools are available.
- Inspects and downloads your uploaded files.
- Pulls frames from the footage and looks at them, along with your photos.
- Checks the voice tools and transcribes your video.
- Reads the full transcript to see what you actually said.
- Prepares a clean voice sample from the recording.
- Chooses the best stretch of footage to work from.
- Clones your voice and generates the new script in it, then checks the result.
- Re-syncs your footage to the new voice, inspects the lip-synced result, zooms in on the face and checks the mouth close up.
Step 6 does more for quality than anything downstream. A voice model built from a stretch with background noise, an overlapping voice or a cough in it carries all of that into every future script, so choosing the sample carefully is worth the time.
What does the result look like?
A short clip of you saying something you never recorded: your face, your voice, your delivery, with the mouth matching the new words. Alongside it sits the voice model, which is the part with lasting value — future scripts do not need new footage.
The clone is only as good as the footage it came from. Well-lit, front-facing, clearly spoken material produces something you would publish; a wobbly clip in a noisy room produces something that looks like what it is.
When does this not work well?
Cloning a person is the one area where the technology's limits and its ethics matter equally, and neither of them is negotiable. Vizard Agent will build what you ask for, so the judgement about whether you should is entirely yours to make.
- Only clone yourself. Someone else's face or voice needs their explicit permission.
- Disclose it where it counts. Some contexts and jurisdictions require you to say so.
- Lip-sync shows at close range. A tight crop reveals more than a mid-shot.
- Poor source footage carries through. Noise in the sample lives in every script.
- Long scripts drift. A minute of clone is more convincing than five.
How do you fix a result that came back wrong?
Say what looks or sounds off. Vizard Agent keeps the voice model, the chosen footage stretch, the transcript and the lip-sync inspections, so the voice or the sync can be redone without starting from your original upload again.
- "The mouth is out." Re-synced and checked at the close-up again.
- "That does not sound like me." A different stretch is used as the sample.
- "It drifts near the end." The script is split into shorter segments.
How does Vizard Agent compare to doing it yourself?
By hand this means recording every script from scratch, which is fine until you have twenty of them and a correction to make in each. The alternative — a generic stock presenter — is not you, which is usually the whole point.
| By hand | Vizard Agent | |
|---|---|---|
| New scripts | Set up and record again | Generated in your cloned voice |
| The sample | Whatever you recorded | The cleanest stretch, chosen deliberately |
| Lip-sync | Notice it on playback | Checked at a mouth close-up |
| Corrections | Re-shoot the whole thing | Re-generate the affected lines |
Common questions
Can Vizard Agent clone anyone? Only with permission. Use Vizard Agent on yourself, or on someone who has explicitly agreed.
How much footage do I need? Not much. Vizard Agent found a usable voice sample within a single uploaded video.
Will it sound like me? Closely, if the sample is clean. Vizard Agent picks the clearest stretch rather than the first one.
Does the mouth match the new words? Yes. Vizard Agent re-syncs the footage and verifies it at a close-up on the mouth.
Can the clone speak another language? Yes, in your voice. That is one of the more useful things to do with a voice model.
Why check the lip-sync zoomed in? Because sync errors that are invisible at full frame become obvious the moment anyone watches on a phone held close, which is how most of this material is actually seen.
Should I tell viewers it is a clone? In many contexts yes, and in some jurisdictions it is required. Check what applies to you.
Can I reuse the voice model later? Yes, and that is the point. Vizard Agent reuses it so new scripts need no new footage.
How long should a clone clip be? Short. Vizard Agent produces a more convincing minute than five minutes.
What footage should I record for this? One clean, well-lit take of yourself talking to camera for a minute or two. Vizard Agent needs a clear voice sample and a steady front-facing stretch far more than it needs interesting content.
Is a clone better than just recording again? Only at volume. For one video, record it. For twenty scripts a month with corrections, Vizard Agent reusing a voice model is the difference between a workflow and a standing commitment.
Does Vizard Agent check the result? Yes. It inspects the sync and the face before delivery.