How to generate a two-shot cinematic scene from character reference images
Give Vizard Agent images of your characters and describe the scene beat by beat. Vizard Agent crops close-ups to study each character properly, generates the first shot, then extracts its final frame and uses it as the opening frame of the second — so the two clips read as one continuous piece rather than as a cut between two guesses.
What is the short version?
Generated video is good for a few seconds at a time and loses the thread beyond that. The way around it is not a longer generation — it is two shots where the second one starts from the exact frame the first one ended on.
- Go to Vizard Agent and give it the character references.
- Describe the scene in order: what happens, then what changes.
- Say where the turn is — the moment the mood shifts.
What do you need before you start?
Character references and a scene with a shape. Vizard Agent studies your references closely before generating anything, and what it needs from you is a scene that actually turns somewhere, because a two-shot structure exists to put the change between the shots.
- The character images. Clear pictures, or frames from a video Vizard Agent can crop from.
- The scene, beat by beat. Where they are, what they do, what changes.
- The turn. The rain stopping, the light breaking, the head lifting. This is where shot one ends.
- The mood. Oppressive to hopeful, tense to calm. It drives the grade and the music.
- The length. Ten seconds across two shots is the natural size for this.
What do you type into Vizard Agent?
Write it like a scene description rather than a prompt. Vizard Agent generates from what you describe, and the detail that pays off most is the physical specifics — the posture, the direction of the gaze, exactly when the weather changes — because those are the things it can render and check.
Prompt
Variants worth knowing:
- References from a video. Vizard Agent can pull the characters out of footage you already have.
- More than two shots. The same chaining works for three or four; each starts from the last frame of the one before.
- Native sound kept. Rain, wind and footsteps come with the generation and can be mixed under music.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real two-shot scene. The pivot is in the middle: rather than generating both shots independently and hoping they match, it hands the end of the first one to the second as its starting point.
- Inspects the uploaded material and extracts frames every few seconds.
- Tiles the frames into a grid and reads it to identify the characters.
- Extracts high-resolution frames to get the character details precisely.
- Crops the character regions and looks at them close up, one at a time.
- Checks the advanced generation tool's options before planning the shots.
- Uploads the character references and generates the first shot.
- Extracts test frames from that shot, tiles them and reviews the composition and quality.
- Takes the last frame of the first shot, uploads it, and generates the second shot from it.
- Reviews the second shot's frames for continuity, joins the two losslessly, and mixes music under the generated rain.
Step 8 is the whole technique. Continuity between generated shots is not something you can ask for in words; handing the second generation the exact frame the first one ended on is how you get it.
What does the result look like?
From the run this page is written from, probed on the delivered file: 720x1280, H.264, 24fps, 10.24 seconds, AAC audio. Vertical, ten seconds, two generated shots joined without a re-encode, with the scene's own rain sound mixed under a chosen music track.
Vizard Agent then ran an analysis over the finished cut to check the motion, the join between the shots and the synchronisation of the picture with the sound.
Ten seconds across two shots is the shape this technique is built for. Each half is short enough to hold its characters steady, and the join between them is where the scene's turn happens.
Expect tens of minutes rather than minutes. This category was not separately measured; comparable work runs a median of 28 to 38 minutes end to end, and across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.
When does this not work well?
Chaining shots holds continuity far better than a single long generation, and it does not make it perfect. Vizard Agent checks each shot's frames before moving on, and generated characters still shift between takes in ways that a close viewer notices.
- Faces drift between shots. Close-ups are where it shows first. Wider framing is more forgiving.
- Long shots wander. Five seconds each is the comfortable range; the drift grows with length.
- Hands and props are the weak point. An umbrella held in shot one may be held differently in shot two.
- The turn has to be visual. "He realises she is sad" is not something generation can show; "she lifts her head" is.
- Characters from a reference are approximations. They resemble the images Vizard Agent studied; they are not copies of them.
- A scene is not a story. Ten seconds carries one turn. Anything with two changes in it needs more shots.
How do you fix a result that came back wrong?
Say which shot. Vizard Agent keeps the character references, both generated shots and the joining frame, so regenerating the second shot leaves the first exactly as it was and still starts from the same frame — which is what keeps a fix from breaking the continuity.
- "Shot two does not look like her." Regenerated from the same reference crops.
- "The rain should stop later." A different turn point in the second shot.
- "Make it feel colder before the change." A grade on the first shot only.
How does Vizard Agent compare to doing it yourself?
By hand this is two prompts to a generation tool, two clips that look like different characters in different weather, and an editor where you discover the join is unusable. There is no way to pass continuity between two separate prompts.
| By hand | Vizard Agent | |
|---|---|---|
| Character consistency | Re-prompt and hope | References cropped and studied |
| Continuity between shots | Impossible to specify | Last frame becomes the next first frame |
| Checking a shot | Watch it | Frames tiled and reviewed |
| Sound | Add it afterwards | Native audio mixed with the score |
Common questions
How many reference images do I need? One clear image per character. Vizard Agent crops close-ups from them, or from a video you upload, to study the details before generating anything.
Can it do more than two shots? Yes. Each new shot starts from the last frame of the previous one, so Vizard Agent can extend the chain as far as the scene needs.
Will the characters look exactly like my references? They resemble them closely. Generation approximates rather than copies, and closer framing makes any difference more visible.
Does the video have sound? Yes. The generated scene carries its own ambience, and Vizard Agent mixes a chosen music track underneath it rather than replacing it.
How long can each shot be? About five seconds each holds up well. Longer generations drift further from the reference.
Can I use frames from my own footage as references? Yes. Vizard Agent extracts and crops them itself rather than needing you to prepare the images beforehand.