How to split a concert recording into one video per song
Upload the show and tell Vizard Agent you want one complete song per video. Vizard Agent transcribes the stage talk, maps the audio energy across each transition at high resolution, cuts test snippets to verify the boundaries by ear, and renders a separate file for every song.
What is the short version?
Finding where a song ends in a live recording is harder than it sounds. The last chord, the applause, the singer thanking the audience and the count-in for the next number all blur together, and cutting on the wrong one leaves a video that starts or finishes badly.
- Go to Vizard Agent and upload the full concert recording.
- Say you want one complete song per video, not highlights.
- Give the song titles if you have a setlist.
What do you need before you start?
The recording and a clear instruction. Vizard Agent finds the boundaries itself, so a setlist is helpful rather than necessary. What matters most is stating that you want whole songs, because the default assumption for a long recording is usually a highlights cut.
- The concert recording. However long the show ran.
- Whole songs, explicitly. Not shorts, not highlights, not excerpts.
- The setlist. For naming the delivered files.
- Where to start and stop. Whether the stage talk belongs to a song or not.
- The shape. Usually unchanged from the source.
What do you type into Vizard Agent?
Rule out the alternatives directly and by name. Vizard Agent reads a line like "not shorts, not highlights, one complete song per video" as a hard constraint, and that single sentence changes the whole method from choosing good moments to finding exact boundaries.
Prompt
Variants worth knowing:
- Stage talk included or cut. Say which; the introduction often belongs to the song.
- Named output files. Give the setlist and each video arrives titled.
- A fade at each end. Rather than a hard cut into applause.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real concert split. Almost all of it is boundary work, and the step that settles it is the one where Vizard Agent cuts short audio snippets around each transition and listens to them properly.
- Probes the recording and transcribes the audio to find the stage talk between songs.
- Examines the exact word timings around every transition.
- Downloads the source locally and extracts the audio for acoustic analysis.
- Calculates the energy profile across each transition to see where the music actually stops.
- Analyses the intros, the endings and the applause as separate acoustic events.
- Extracts frames from each transition and inspects them visually.
- Maps the sound energy at high resolution across all four transitions.
- Cuts short audio snippets at each candidate boundary and checks them individually before committing.
- Renders one video per song, then verifies the first and last frames of every file and checks the audio levels at each entry and exit.
Step 8 is what makes the boundaries right rather than approximately right. An energy curve tells you where something changed; only listening to the moment tells you whether that change was the song ending or the drummer counting in.
What does the result look like?
From the run this page is written from, probed on one delivered file: 1920x1080, H.264, 30fps, 125.53 seconds, AAC audio. Widescreen, just over two minutes for that song, delivered as one of three separate files, each starting and ending on a verified boundary rather than a guess.
Delivering separate files rather than one video with chapters is what the format is for. Each song becomes something you can post, title and share on its own, which is the whole reason to split a show in the first place.
When does this not work well?
A live recording is one continuous performance that happens to contain songs, and not every show divides cleanly along those lines. Vizard Agent finds the boundaries acoustically wherever they exist, and some music genuinely does not have them at all.
- Medleys and segues have no gap. Where one song runs into the next, any cut is arbitrary.
- The performance rights are yours. Splitting a recording does not change who owns the songs.
- Long stage talk is ambiguous. Whether a three-minute story belongs to the next song is your call.
- A single fixed camera limits it. Each file inherits whatever framing the show was shot with.
- Crowd noise masks endings. Loud rooms make the acoustic boundary less distinct.
How do you fix a result that came back wrong?
Name the song and the end that is wrong. Vizard Agent keeps the transcript, the energy maps, the tested snippets and the verification frames, so moving a boundary is a re-render of two files rather than a fresh analysis of the whole show.
- "Song two starts too late." Re-cut from the mapped transition.
- "Keep the introduction with the song." The boundary moves back before the stage talk.
- "It ends on applause." Trimmed to the last note and re-verified.
How does Vizard Agent compare to doing it yourself?
By hand this means scrubbing a two-hour recording with a waveform open, guessing at each boundary because applause looks like music on a waveform, and exporting one file at a time. The mistakes only show up later, when a video opens on the last four seconds of the previous song.
| By hand | Vizard Agent | |
|---|---|---|
| Finding boundaries | Guess from the waveform | Energy mapped at high resolution per transition |
| Confirming them | Assume the waveform is right | Snippets cut and checked at each candidate |
| Stage talk | Cut it or leave it | Transcribed, so it can be assigned deliberately |
| Checking the output | Watch the start of each | First and last frames verified on every file |
Common questions
Does this work for other long recordings? Yes. A lecture series, a conference day or a DJ set all split the same way. Vizard Agent looks for the acoustic and spoken markers that separate one section from the next, then verifies each boundary by cutting and checking it rather than trusting the curve.
Do I need a setlist? No. Vizard Agent finds the boundaries acoustically, and a setlist helps it name the files.
Will each video be a complete song? Yes, that is the format. Vizard Agent verifies the first and last frames of every file.
What about the talking between songs? Vizard Agent transcribes it, so you can decide whether it belongs to a song or gets cut.
Can Vizard Agent handle a two-hour show? Yes. It works from the audio energy profile rather than watching the whole thing.
Why not just cut on the applause? Because applause starts before the last chord decays and often continues into the next introduction. Cutting there clips the ending of one song and puts the crowd noise at the front of the next one.
Does Vizard Agent check the audio levels? Yes, at the entry and exit of every file, so no video opens or closes on a jump in volume.
Can Vizard Agent name the files? Yes, from the setlist you supply.
What if two songs run together? A segue has no real boundary. Vizard Agent will tell you where the closest one is and the choice is yours.