Vizard Agent

How to cut a video sales letter from footage you already have

Last updated 2026-08-25 · 7 min read

Upload your raw videos and tell Vizard Agent you want a video sales letter. Vizard Agent transcribes the footage, finds the sections that carry the argument, marks the frame with a coordinate grid before placing any text, and cuts a VSL that moves through problem, proof and offer.

What is the short version?

A video sales letter is a structure, not a style. It has to open on a problem the viewer recognises, prove the claim, then ask for something, and the footage you already shot usually contains all three parts in the wrong order.

  1. Go to Vizard Agent and upload your raw videos.
  2. Say it is a VSL and what it is selling.
  3. Say whose voice narrates it.

What do you need before you start?

Raw footage and a clear offer. Vizard Agent reads what was said across all of your clips and finds the sections that carry each part of the argument, so a written script is unnecessary. What you do need is a decision about what you are asking the viewer to do at the end.

What do you type into Vizard Agent?

Name the format explicitly rather than describing what you want the video to feel like. Both "VSL" and "video sales letter" tell Vizard Agent to build toward a call to action rather than toward a highlight reel, and that single word changes which parts of your footage get selected.

Prompt

Cut a [VSL] from these [4] videos for [what you sell]. Open on the problem, prove it, then the offer. Narrate with [my cloned voice].

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real VSL cut from four raw videos. The step people do not expect is the coordinate grid: before placing a single word on screen, it marks the frame so text lands where it is meant to.

  1. Reads the workspace and downloads your uploaded videos.
  2. Transcribes the footage and reads what was actually said across all of it.
  3. Checks the available voices, music and caption styles before committing to any.
  4. Generates the narration, using a cloned voice for the scene that needs to sound like you.
  5. Extracts frames and marks them with a coordinate grid, then reads the grid to plan the on-screen layout.
  6. Builds panels of the candidate B-roll and reviews them by eye to choose the shots.
  7. Measures the loudness of the narration and pulls word-level timings for the cut.
  8. Cuts the scenes against those timings, adds captions, and mixes music under the voice.
  9. Watches the assembled VSL back, checks the tone transition holds, and fixes what does not land.

Step 5 is what stops text from sitting over a face or drifting off the safe area. Reading the actual pixel coordinates of the frame is slower than guessing, and it is the reason the layout survives the first playback.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 54.47 seconds, AAC audio. Vertical, just under a minute, four raw videos reduced to one argument, narrated in the creator's own cloned voice with captions and music underneath.

Fifty-four seconds is a short VSL, which is the right length when it is running as an ad. The same method produces a five-minute version; only the amount of proof in the middle changes.

When does this not work well?

A VSL asks the viewer for something, and that makes it a harder format to get right than a highlight reel. Some of the limits Vizard Agent runs into are in the footage itself, and some are in the claim you are making.

How do you fix a result that came back wrong?

Say which part of the argument is failing. Vizard Agent holds the transcript, the narration timings, the B-roll panels and the coordinate map of the frame, so it can restructure the argument without recutting the whole thing from scratch.

How does Vizard Agent compare to doing it yourself?

By hand a VSL means scrubbing four videos for the lines that make the case, writing narration around them, recording it, then laying out text that does not collide with what is in frame. The layout alone usually takes a couple of passes.

By hand Vizard Agent
Finding the argument Scrub and remember Read from the full transcript
Narration Record it yourself Generated or in your cloned voice
Text placement Eyeball it, adjust Placed against a real coordinate grid
B-roll choice Guess from filenames Chosen from panels, by eye

Common questions

What is a VSL? A video sales letter: problem, proof, then offer. Vizard Agent builds toward the ask rather than toward a montage.

How many videos do I need? Four worked on this run. Vizard Agent can work from fewer if the argument is all in one of them.

Can Vizard Agent use my own voice? Yes, cloned, for a scene that was never filmed. Only clone a voice you have the right to use.

Do I need to write the script? No. Vizard Agent reads what you already said on camera and builds the narration around it.

How long should a VSL be? Under a minute as an ad, several minutes on a landing page. Say which and Vizard Agent adapts the structure.

Why does Vizard Agent map the frame coordinates? Because text placed by guesswork lands on faces and drifts outside the safe area. Reading the actual pixel positions first means the layout survives the first playback instead of needing two correction passes.

Can I see the B-roll before it is used? Yes. Vizard Agent builds panels of every candidate shot and chooses from them visually.

Does the tone change through the video? It can. Ask for a transition and Vizard Agent shifts the music, the pacing and the captions with it.

Will it make claims about my product? Vizard Agent works only from what your footage says. Whatever you assert remains your responsibility.

Can I get a longer cut too? Yes. Vizard Agent produces a longer version with more proof from the same transcript and selections.