How to break a reference video into its actual scenes
Ask Vizard Agent to cut the reference into its real scenes and give them to you as separate files. The boundaries that matter are the ones the video already uses — the fades it passes through, where its sentences end, the pause before the music drops — and equal intervals will not find any of them.
What is the short version?
You want to make something in the same style as a video you have. Describing that style in words loses most of it. Vizard Agent takes the reference apart on its own structural boundaries, so you end up with its scenes rather than an impression of them.
- Give Vizard Agent the reference and ask it to be broken into scenes.
- Let it find the boundaries from the fades and the speech.
- Look at the scenes as files before you build anything.
What do you need before you start?
The reference video itself, as a file or a link it can reach. An analysis of something it cannot open is an analysis of nothing, and a decomposition is the one job where there is no way to work around a missing source.
It also helps to say what you intend to do with it. A structural breakdown for copying pacing is a different cut from one for copying the graphics, and saying which changes what gets measured.
What do you type into Vizard Agent?
Ask for the pieces, not a summary. The instinct is to ask what the style is; the useful request is to have the thing divided up, because a list of scenes with real durations tells you more than any adjective would.
Analyse this in as much detail as possible. Break it into its actual scenes and give me each one as a separate file I can watch, with its duration. Cut on the boundaries the video itself uses — its fades to black and where the sentences end — not on fixed intervals. I want to rebuild in the same style and format.
The last clause matters because it tells Vizard Agent what precision is for. A decomposition you intend to build from needs exact durations; one you are only curious about does not.
What does Vizard Agent actually do?
Vizard Agent measures the reference from several directions before cutting anything, because structure is carried by different things in different videos — sometimes the edit, sometimes the speech, sometimes the music itself. Only then does it cut, and it cuts twice, at two different grains.
In the session behind this article it:
- Built an edit map, a full transcript and a contact sheet of the reference.
- Checked the motion, the transitions, the sync and how the audio was delivered.
- Examined the title animation frame by frame, and where the music paused.
- Went through the first and last twenty-three seconds frame by frame.
- Looked at the waveform's shape and density, and checked the pause and drop around thirty seconds.
- Cut the video into twenty-two scenes, verified each duration, and trimmed the audio tails.
- Delivered all twenty-two as separate files with links to watch them.
- Re-cut on exact spoken-phrase boundaries using word timings.
- Located the exact fade-to-black intervals and cut eleven larger scenes on those.
The two passes are the point of the exercise. Phrase boundaries give you the fine grain of how the thing is written, while the blackouts give you its sections and how long each one runs. Either one on its own tells you about half the structure, and the half it leaves out is usually the half you needed.
What does the result look like?
A folder of scenes with real durations, and a second, coarser set of sections. You can watch the third scene on its own and see exactly how long the reference spends on a beat like that — which is the number you actually needed.
The durations are what transfer to your own work. Knowing a reference gives four seconds to its opening title and eleven to its first section is a specification you can build against; knowing that it feels "fast-paced" is not, and Vizard Agent gives you the former.
When does this not work well?
When the reference has no structural devices to find. A continuous single-take video with no fades and no clear sections does not decompose at all, and a cut placed on arbitrary boundaries would invent a structure that was never in the original.
It also misleads if you copy the decomposition too literally. A reference's scene lengths suit its own content, and pouring different material into the same durations usually produces something that feels borrowed rather than similar.
And a reference you do not have rights to is a reference for structure only. Timings, pacing and shot order are fair to learn from; its footage, music and graphics are not yours to reuse.
How do you fix a result that came back wrong?
If the scenes are too fine or too coarse, name the device you want them cut on. "Cut on the fades only" and "cut on every shot change" are different requests, and Vizard Agent can do either from measurements it already has.
If a boundary lands mid-sentence, say so. It re-cuts from word timings rather than nudging the timecode, which fixes the cause rather than the instance.
If you wanted the graphics rather than the structure, ask for that specifically. Vizard Agent can go frame by frame through a title animation, which is a different measurement from scene boundaries.
How does Vizard Agent compare to doing it yourself?
Taking a reference apart by hand means scrubbing for the cuts and writing down timecodes, and the tedium is why most people skip it and work from memory instead. The result of skipping it is a video that resembles the reference in tone and nothing else.
The habit worth copying is cutting on the original's own devices rather than your own instincts. Its fades mark where its editor thought a section ended, and that is better information than any judgement you would make about it from outside.
Common questions
Why not just describe the style? Descriptions lose the durations, and Vizard Agent works from durations.
How many scenes will I get? As many as the video has. Vizard Agent found twenty-two, then eleven sections.
Can I watch them individually? Yes. Vizard Agent delivers each scene as its own file.
What boundaries does it use? Vizard Agent uses fades to black and spoken-phrase ends from word timings.
Why two passes? Phrases give Vizard Agent the fine grain; blackouts give the sections.
Can it check the title animation too? Yes. Vizard Agent goes through it frame by frame if asked.
Does it look at the music? Yes. Vizard Agent notes where it pauses and where it drops.
What if there are no fades? Then say which device to cut on. Vizard Agent will not invent one.
Can I reuse the reference's footage? No. Have Vizard Agent learn its structure, not its material.
Are the durations exact? Yes, and Vizard Agent verifies each one after cutting.
Should I copy the lengths directly? Not blindly. They suit the reference's own content, not yours.
Can it do this for several references? Yes, and comparing them is more useful than one alone.
What about the audio tails? Vizard Agent trims them so each scene ends cleanly.
Can I get the transcript as well? Yes. Ask Vizard Agent for it alongside the scenes.