How to cut a video sales letter from footage you already have
Upload your raw videos and tell Vizard Agent you want a video sales letter. Vizard Agent transcribes the footage, finds the sections that carry the argument, marks the frame with a coordinate grid before placing any text, and cuts a VSL that moves through problem, proof and offer.
What is the short version?
A video sales letter is a structure, not a style. It has to open on a problem the viewer recognises, prove the claim, then ask for something, and the footage you already shot usually contains all three parts in the wrong order.
- Go to Vizard Agent and upload your raw videos.
- Say it is a VSL and what it is selling.
- Say whose voice narrates it.
What do you need before you start?
Raw footage and a clear offer. Vizard Agent reads what was said across all of your clips and finds the sections that carry each part of the argument, so a written script is unnecessary. What you do need is a decision about what you are asking the viewer to do at the end.
- Your videos. Four is enough; more gives more choice.
- The offer. What you want the viewer to do.
- The audience. Who this is aimed at.
- The voice. Cloned from you, or generated.
- The length. VSLs run short or very long; pick one.
What do you type into Vizard Agent?
Name the format explicitly rather than describing what you want the video to feel like. Both "VSL" and "video sales letter" tell Vizard Agent to build toward a call to action rather than toward a highlight reel, and that single word changes which parts of your footage get selected.
Prompt
Variants worth knowing:
- Tone transition. Sober at the start, warmer at the offer.
- A new scene in your voice. For a line that was never filmed.
- Panels of the B-roll. So you can see the candidate shots before they are used.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real VSL cut from four raw videos. The step people do not expect is the coordinate grid: before placing a single word on screen, it marks the frame so text lands where it is meant to.
- Reads the workspace and downloads your uploaded videos.
- Transcribes the footage and reads what was actually said across all of it.
- Checks the available voices, music and caption styles before committing to any.
- Generates the narration, using a cloned voice for the scene that needs to sound like you.
- Extracts frames and marks them with a coordinate grid, then reads the grid to plan the on-screen layout.
- Builds panels of the candidate B-roll and reviews them by eye to choose the shots.
- Measures the loudness of the narration and pulls word-level timings for the cut.
- Cuts the scenes against those timings, adds captions, and mixes music under the voice.
- Watches the assembled VSL back, checks the tone transition holds, and fixes what does not land.
Step 5 is what stops text from sitting over a face or drifting off the safe area. Reading the actual pixel coordinates of the frame is slower than guessing, and it is the reason the layout survives the first playback.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 54.47 seconds, AAC audio. Vertical, just under a minute, four raw videos reduced to one argument, narrated in the creator's own cloned voice with captions and music underneath.
Fifty-four seconds is a short VSL, which is the right length when it is running as an ad. The same method produces a five-minute version; only the amount of proof in the middle changes.
When does this not work well?
A VSL asks the viewer for something, and that makes it a harder format to get right than a highlight reel. Some of the limits Vizard Agent runs into are in the footage itself, and some are in the claim you are making.
- A weak offer stays weak. Editing cannot make an unclear proposition clear.
- Proof has to exist in the footage. If you never said it, it is not there.
- Cloned voices need consent. Only clone a voice you are entitled to use.
- Claims are yours. Anything you assert about results is your responsibility, not the edit's.
- Very short VSLs lose the middle. Under a minute, proof gets compressed hard.
How do you fix a result that came back wrong?
Say which part of the argument is failing. Vizard Agent holds the transcript, the narration timings, the B-roll panels and the coordinate map of the frame, so it can restructure the argument without recutting the whole thing from scratch.
- "The hook is too slow." Re-selected from the transcript against the opening.
- "Show the product earlier." Re-ordered from the B-roll panels already built.
- "The offer is buried." The final scene is rebuilt and re-timed.
How does Vizard Agent compare to doing it yourself?
By hand a VSL means scrubbing four videos for the lines that make the case, writing narration around them, recording it, then laying out text that does not collide with what is in frame. The layout alone usually takes a couple of passes.
| By hand | Vizard Agent | |
|---|---|---|
| Finding the argument | Scrub and remember | Read from the full transcript |
| Narration | Record it yourself | Generated or in your cloned voice |
| Text placement | Eyeball it, adjust | Placed against a real coordinate grid |
| B-roll choice | Guess from filenames | Chosen from panels, by eye |
Common questions
What is a VSL? A video sales letter: problem, proof, then offer. Vizard Agent builds toward the ask rather than toward a montage.
How many videos do I need? Four worked on this run. Vizard Agent can work from fewer if the argument is all in one of them.
Can Vizard Agent use my own voice? Yes, cloned, for a scene that was never filmed. Only clone a voice you have the right to use.
Do I need to write the script? No. Vizard Agent reads what you already said on camera and builds the narration around it.
How long should a VSL be? Under a minute as an ad, several minutes on a landing page. Say which and Vizard Agent adapts the structure.
Why does Vizard Agent map the frame coordinates? Because text placed by guesswork lands on faces and drifts outside the safe area. Reading the actual pixel positions first means the layout survives the first playback instead of needing two correction passes.
Can I see the B-roll before it is used? Yes. Vizard Agent builds panels of every candidate shot and chooses from them visually.
Does the tone change through the video? It can. Ask for a transition and Vizard Agent shifts the music, the pacing and the captions with it.
Will it make claims about my product? Vizard Agent works only from what your footage says. Whatever you assert remains your responsibility.
Can I get a longer cut too? Yes. Vizard Agent produces a longer version with more proof from the same transcript and selections.