How to split a recording into segments and drop the retakes
Upload the take and tell Vizard Agent how many segments you want and that you fluffed lines along the way. Vizard Agent reads the transcript for the places where a sentence starts over, confirms each restart against the silence around it, and cuts the segments with the bad takes removed.
What is the short version?
Recording ten short videos in one sitting is far easier than recording them separately, and it leaves you with one long file containing ten topics and every mistake you made. Splitting it means finding both the topic boundaries and the restarts, which are different problems.
- Go to Vizard Agent and upload the recording.
- Say how many segments you want and roughly how long each should be.
- Say that you made speaking errors and started again.
What do you need before you start?
The recording and a count. Vizard Agent finds the topic boundaries and the retakes from the transcript, so you do not need to note where anything happened. Telling it that retakes exist is the instruction that matters, because otherwise a stumble followed by a restart reads as two attempts at the same sentence and both survive.
- The recording. One long take, at whatever resolution you shot.
- How many segments. Ten is a common target.
- The length each. Sixty seconds is a natural unit.
- That retakes exist. Say so explicitly.
- The topics, if you know them. Vizard Agent finds them either way.
What do you type into Vizard Agent?
Say both parts in one instruction. Vizard Agent handles the segmenting and the retake removal as a single job, and mentioning the errors up front is what stops a false start being treated as the beginning of a new segment.
Prompt
Variants worth knowing:
- Keep the original resolution. Say so if you shot in 4K and want it back.
- Even lengths, or natural ones. Sixty seconds each, or wherever each topic ends.
- A named file per segment. If you already know the topics.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real segmenting job. Almost the entire process is reading, and the two things it is looking for at once are where a topic begins and where a sentence was abandoned.
- Transcribes the recording and inspects the transcript segments.
- Reads the full transcript to identify the distinct sections and the retakes.
- Structures the transcript into numbered lines and reads all of them to map the topics.
- Reviews every line specifically for section boundaries and speech errors.
- Finds all the candidate segment headers and restart points.
- Checks the word timings and extracts precise timings for each proposed section.
- Detects the long pauses and transitions between sections.
- Confirms each error restart point and checks the verbatim speech in every section.
- Detects the exact silence points between takes, verifies every cut boundary, benchmarks the render speed, then cuts all ten segments with the retakes removed and reviews them as an inspection grid.
Step 9's first half is what makes the cuts clean. A restart usually has a short silence in front of it, and using that silence as the boundary means the bad take disappears without clipping the first word of the good one.
What does the result look like?
From the run this page is written from, probed on one delivered segment: 3840x2160, H.264, 29.97fps, 64.06 seconds, AAC audio. Full 4K, just over a minute, one of ten separate files, each starting on a new topic with the false starts cut out.
Keeping the original 4K rather than delivering a proxy matters here because these are finished videos, not intermediates. Vizard Agent benchmarked the encode first specifically so ten 4K exports stayed practical.
When does this not work well?
Detecting a retake means inferring intent from a transcript, and people stop and start a sentence again for several different reasons. Vizard Agent finds the clear cases reliably and flags what it cut, and the ambiguous ones are worth a look from you before publishing.
- Not every repetition is a mistake. Saying something twice for emphasis looks like a retake in text.
- Half-corrections are hard. Changing one word mid-sentence is not the same as starting over.
- A rambling take has no boundaries. If the topics blur into each other, so will the segments.
- Ten segments needs ten topics. Vizard Agent will not invent divisions that are not there.
- 4K exports take time. Ten of them is a real render job rather than an instant one.
How do you fix a result that came back wrong?
Name the segment. Vizard Agent keeps the numbered transcript, the section map, the restart points and the verified boundaries, so restoring a take or moving a split is a re-render of one file rather than a fresh pass through the recording.
- "That was not a mistake, keep it." Restored from the transcript and re-cut.
- "Segment six starts mid-thought." Moved back to the mapped section boundary.
- "Split segment three in two." Divided at the detected pause between its topics.
How does Vizard Agent compare to doing it yourself?
By hand this means watching your own take back with a notepad, marking ten boundaries and every fluff, then exporting ten files one at a time. The retakes are what make it slow, because each one needs the cut placed in the silence rather than on the waveform, and there are usually more of them than you remember.
| By hand | Vizard Agent | |
|---|---|---|
| Finding topic boundaries | Watch and mark | Mapped from a numbered transcript |
| Spotting retakes | Remember them | Restart points found and confirmed |
| Placing the cut | On the waveform | On the detected silence before the restart |
| Ten exports | One at a time | Encode benchmarked, all ten rendered |
Common questions
Does this work for other batch recordings? Yes. A run of product answers, a set of lesson intros or a morning of client updates all split the same way. Vizard Agent maps the topics from the transcript and removes the false starts wherever they occur.
How many segments can Vizard Agent produce? Ten on this run. It works from the topics in the recording rather than a fixed limit.
How does Vizard Agent know I made a mistake? It reads the transcript for sentences that stop and start again, then confirms each one against the silence around it.
Will Vizard Agent keep my 4K? Yes. It benchmarked the encode first so ten full-resolution exports stayed practical.
What if a segment runs long? Vizard Agent calculates each section's duration and will tell you where a topic exceeds your target.
Why cut on the silence rather than the waveform? Because the gap before a restart is where the sentence genuinely ended. Cutting there removes the bad take without clipping the breath or the first syllable of the good one, which is what makes the join inaudible.
Can Vizard Agent name the files? Yes, if you give it the topics.
Does Vizard Agent check the finished segments? Yes. It verifies every duration and reviews all ten as a visual inspection grid.
Can I record several topics in one sitting? That is what this format is for. One take, ten videos, retakes removed.