How to recreate commentary your microphone never recorded
Send Vizard Agent your footage plus timestamped notes on what you were saying, and one rule: fill small gaps from what is visible on screen, but invent nothing important. It picks one voice, writes the lines against the actual moments, and lands each reaction on the frame it belongs to.
What is the short version?
You played for an hour, something genuinely funny happened, and the microphone was never armed. The footage is fine and the commentary is gone. Vizard Agent rebuilds the commentary from your notes rather than writing a new one.
- Write Vizard Agent notes with timestamps.
- Say what it may and may not fill in.
- Check each line lands on the moment it reacts to.
What do you need before you start?
The recording and your notes. The notes are the whole job — the more specific they are, the more it sounds like you, and a note that says "6:24 — this is where he starts camping me and I lost it" is worth ten minutes of generic reaction writing.
You also want a decision about the voice. A commentary track needs to sound like one person across the whole video, so the choice gets made once and used for every line. Say what it should sound like, including any effect you want on it, because that is easier to set up front than to apply to fifteen finished lines.
What do you type into Vizard Agent?
Give the notes, the voice and the boundary. The boundary is the part people leave out and the part that protects you — commentary is opinion, and an invented opinion published in your voice is a real problem rather than a stylistic one.
My mic did not record. Here are my notes with timestamps. Recreate the commentary in one voice. Follow my notes closely, you can fill small gaps from what is obviously happening, but do not invent events or opinions I did not have.
That was close to the real brief behind this article, and Vizard Agent treated the last clause as a hard limit — asking rather than guessing when an important beat was not covered by the notes.
What does Vizard Agent actually do?
It finds the moments before it writes the lines. A reaction that arrives half a second after the thing it reacts to reads as dubbing, so each note gets resolved to an actual span in the footage first, and the line is written and timed against that span.
In that session the working order was:
- Surveyed the whole recording to find where the story actually escalates.
- Read the notes section by section against what the footage shows at those points.
- Pinned specific moments — an escape at about 1:42, a repeat encounter at 3:10, a return thirteen minutes in — to exact spans rather than rough minutes.
- Auditioned voices and chose one, then generated all fifteen commentary beats with it.
- Applied the requested pitch and echo treatment to the voice, then checked every processed line's duration and loudness.
- Built the cut around the commentary, and measured the music against the speech so the voice dominates whenever it speaks.
Fifteen lines across an hour is deliberately sparse. Commentary that narrates what the viewer can already see is the commonest way this goes wrong, and the notes are what keep it to the moments that actually need a voice.
What does the result look like?
A cut with commentary that sounds authored rather than generated — one voice, consistent treatment, reacting to specific things at the moments they actually happen. The lines are yours in content and in timing, even though the voice that delivers them is not.
The mix is measured rather than eyeballed. Vizard Agent checks the voice against the music over the exact spoken windows rather than over a longer stretch that includes the silence afterwards, which is the difference between a mix that reads as balanced and one that only measures as balanced.
When does this not work well?
When the notes are thin. "Funny bit around here" gives Vizard Agent nothing to be faithful to, and the honest options are to write a line that is generic or to leave the moment silent. It will tell you which moments your notes do not cover rather than filling them in.
It is also not a substitute for a real reaction on a channel where your voice is the draw. A recreated line is accurate; it does not have the timing of something said in the moment. For a one-off rescue that is fine, and as a habit it is not.
And it cannot reconstruct what you said if you do not remember. The notes are a memory exercise done afterwards, and an hour of play is a lot to reconstruct. Expect to cover the beats that mattered, not the whole session.
How do you fix a result that came back wrong?
If a line does not sound like you, rewrite that line yourself and send it back. Vizard Agent regenerates it in the same voice with the same treatment, so one line can change without the rest of the track losing its consistency.
If a reaction lands late, give the timestamp of what it should be reacting to. That is a timing fix rather than a rewrite, and Vizard Agent re-pins the line to the correct span.
If the voice is wrong overall, say so early. Changing the voice means regenerating every line, which is cheap at line five and annoying at line fifteen — which is why the choice is made once, before any of them are generated.
How does Vizard Agent compare to doing it yourself?
You could re-record it yourself, watching the footage back and talking over it. That is the better-sounding option and most people who try it discover that reacting to something you already know the outcome of is much harder than reacting live.
The alternative by hand is writing the script, generating the lines, and then spending the real time on placing them — finding the exact frame each reaction belongs to and ducking the music under each one. Vizard Agent does the placement against measured spans and the ducking against measured levels, which is the part that takes an evening and produces no creative satisfaction.
Common questions
Do I have to write out every line? No. Notes are enough, and Vizard Agent writes the lines from them.
Can it use my own voice? Only with a voice sample you provide and the right to use it. Otherwise Vizard Agent picks a consistent voice.
How many lines is right for an hour? Fewer than you think. Fifteen was right in the session behind this article.
Will it narrate what I can already see? Not unless you ask. Vizard Agent keeps commentary to reactions rather than description.
What if my notes contradict the footage? Vizard Agent says so rather than choosing. The footage is the check on the notes.
Can I add an effect to the voice? Yes. Say what it should sound like and Vizard Agent applies it consistently to every line.
Will the music drown it? No. Vizard Agent measures the voice against the music over the spoken windows and ducks to a real target.
Can I get captions from it? Yes, and they will be accurate, because Vizard Agent transcribes the generated speech for word timings.
What if I remember more later? Send the extra notes. Vizard Agent adds lines without regenerating the existing ones.
Does the game audio stay? Yes, underneath. Vizard Agent checks whether it is usable before deciding how much to keep.
Can it match a specific creator's style? It can match a described style. Vizard Agent will not imitate a named person's voice.
What about swearing? It follows your notes. Vizard Agent does not add profanity you did not write.
How long does this take? Most of the time goes on finding the moments, not on generating the voice.
Can I approve the script first? Yes. Ask Vizard Agent for the written lines before any audio is generated.