How to structure a comparison video that opens on a question
Write the opening question and describe what each side should show. Vizard Agent records the narration, generates a shot for every beat of both sequences, animates them, builds the on-screen graphics, and generates an extra shot when the argument needs a counterpart the original brief did not include.
What is the short version?
A comparison video works by asking something before it explains anything at all, which is why the structure matters more than the footage. The question buys you the first three seconds, the two sequences do the actual arguing, and the ending has to resolve the comparison rather than simply stopping.
- Go to Vizard Agent and write the question the video opens on.
- Describe each sequence — what it shows and what it proves.
- Say the length and the language.
What do you need before you start?
A question and two sides. Vizard Agent generates every shot, records the narration and builds the graphics, so no footage is needed at all. The question is doing the most work in the whole brief, because it sets what both sequences have to be comparable on.
- The opening question. Written out, as it will appear.
- Both sequences. What each one shows, beat by beat.
- The resolution. What the viewer should conclude.
- The length. Around a minute for two sequences plus a close.
- The language. For the narration and the captions.
What do you type into Vizard Agent?
Write it as numbered sequences. Vizard Agent generates one shot per beat, so a brief laid out as sequence one, sequence two and a closing maps directly onto the shots it builds and the graphics it places over them.
Prompt
Variants worth knowing:
- Animated on-screen graphics. They carry the argument for muted viewing.
- A positive counterpart shot. So the ending resolves rather than trailing off.
- High-contrast captions. Comparison videos are watched without sound.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real comparison video. Two corrections shaped the result: the graphics came back with type far too small, and the ending needed a shot the original brief never asked for.
- Loads the guidance for this format and checks the generation and caption tools.
- Searches for a voice and music, then generates the narration and the first images.
- Measures the narration's duration and generates all fifteen images plus the music bed.
- Assembles a contact sheet and reviews every generated image before animating anything.
- Builds the full voice track and animates all fifteen shots.
- Reviews the animated shots in three groups and transcribes the narration for captions.
- Creates the animated on-screen graphics, checks them, and rebuilds them with much larger type.
- Assembles the picture, composites the graphics and measures the voice and music levels.
- Generates a renovated-house shot as a counterpart to the neglected one, compares the two, replaces the final beat, then rebuilds the captions with better contrast and re-renders.
Step 9 is the one worth copying. The brief described the problem twice and never described the fix, and a comparison video that ends on the second bad example leaves the viewer without the conclusion the whole structure was building towards.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 67.73 seconds, AAC audio. Vertical, sixty-eight seconds, fifteen generated and animated shots across both sequences, narrated, with animated graphics and high-contrast captions.
Sixty-eight seconds is what two sequences and a resolution actually need. Vizard Agent gives each side enough beats to be recognisable before the comparison lands, which a thirty-second cut would not allow.
When does this not work well?
The comparison at the centre of this format is an argument, and Vizard Agent builds the argument you write rather than testing whether it actually holds up. Everything on screen is generated illustration, which suits a rhetorical structure well and rules out anything meant as evidence.
- The comparison has to be fair. A rigged pairing reads as one to anyone paying attention.
- Generated shots illustrate, they do not evidence. Neither side is a photograph of anything real.
- A weak question sinks it. If the opening does not make someone curious, nothing after it matters.
- Two sides is the format. A three-way comparison needs a different structure entirely.
- Claims are yours. Vizard Agent narrates what you write and does not check it.
How do you fix a result that came back wrong?
Name the sequence or the beat. Vizard Agent keeps all fifteen generated shots, the narration timings, both graphic versions and the animated clips, so replacing a shot or rewriting a line is a re-render rather than a fresh generation round.
- "The ending does not resolve." A counterpart shot generated and swapped in, as here.
- "The graphics are unreadable on a phone." Rebuilt with larger type.
- "That side is not represented fairly." Regenerated to a corrected description.
How does Vizard Agent compare to doing it yourself?
By hand this means storyboarding fifteen shots, sourcing or generating each one, and building the graphics in a design tool at whatever size looks right on a monitor. On a phone that type turns out to be half the size it needed to be, and by then the whole video is assembled around it.
| By hand | Vizard Agent | |
|---|---|---|
| The shots | Source or generate one at a time | Fifteen generated, reviewed, then animated |
| Graphic type size | Judged on a monitor | Checked on frame and rebuilt much larger |
| A missing beat | Ship without it | Generated when the structure needed it |
| Caption contrast | One setting | Rebuilt after reviewing the render |
Common questions
Does this structure work for other subjects? Yes, and it is one of the most transferable formats there is. Anything with a neglected version and a maintained one — a business, a skill, a relationship, a piece of equipment — fits the same four-sequence shape that Vizard Agent builds here.
Do I need any footage? None. Vizard Agent generates every shot in both sequences.
How long should a comparison video be? Around a minute. Each side needs several beats before the comparison registers.
Will Vizard Agent write the narration? Yes, from your sequence descriptions, in the language you name.
Can Vizard Agent add on-screen graphics? Yes, animated, and it rebuilt them here once it saw the type was too small.
Why does the opening question matter so much? Because it is the only part of the video that has to work before anyone has invested anything. If the question does not make someone want the answer, the two sequences never get watched.
What if my brief has no ending? Vizard Agent will tell you, and it generated the missing counterpart shot itself on this run.
Can Vizard Agent do this in my language? Yes, for the narration, the captions and the on-screen text.
Does Vizard Agent check the balance of the audio? Yes. It measures the voice against the music before rendering.
Can I compare three things? Not in this structure. Three-way comparisons need a different format.
Does Vizard Agent check the finished video? Yes. It reviews the render, rebuilds what fails and has the final cut analysed.