Vizard Agent

How to cut the dead space out of a clothing showcase

Last updated 2026-08-24 · 7 min read

Upload the take and tell Vizard Agent to cut everything that is not the product. Vizard Agent transcribes the speech, splits the take into shots, zooms into the garment details to see what is actually being shown, and rebuilds the video with the dead space removed and the captions placed clear of the frame's existing graphics.

What is the short version?

A product take recorded on a phone is mostly reaching, adjusting and finding the angle, with maybe eight usable seconds inside it. Cutting the dead space is the whole edit, and the rest is making those eight seconds look deliberate.

  1. Go to Vizard Agent and upload the take.
  2. Say to remove the dead space and keep what shows the product.
  3. Say the register and the language for the captions.

What do you need before you start?

The take and the product names. Vizard Agent finds the shot boundaries and the usable stretches itself, so nothing needs marking. Telling it what the items are is worth doing, because the on-screen text names them and a transcription of casual speech will not always get them right.

What do you type into Vizard Agent?

Say what to remove rather than what to keep. Vizard Agent reads "cut the extra parts" as an instruction to find every stretch where nothing is being shown, which is a more reliable brief than trying to list the moments you want.

Prompt

I am showing [the items] in this video. Edit it properly, cut out the extra parts, and give me something an audience will actually watch. Captions in [language].

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real product showcase. The step people skip is the seventh: it zoomed into the shirt's print and the trousers' fabric to see what the video was actually showing before deciding which seconds to keep.

  1. Probes the video and extracts the speech transcript.
  2. Extracts frames and analyses the footage and the speech together.
  3. Reviews the audio track to see where the useful speech sits.
  4. Splits the take into shots and builds an overview of every cut.
  5. Reviews all the frames at once to see the whole take in one view.
  6. Searches for trendy music and analyses each candidate's structure and rhythm.
  7. Zooms into the garment details — the shirt's print, the trousers' fabric — to see what is being shown.
  8. Checks the first and last frames of every shot and reviews the sequence for continuity.
  9. Sources a typeface for the language and tests it on frame, cuts and joins the usable sections, then repositions the captions after finding them colliding with stickers already in the picture.

Step 9's ending is specific to this format. Product videos filmed for social often already have stickers or text burned in from the phone app, and captions added on top will land on them unless somebody checks the actual frames.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 12.97 seconds, AAC audio. Vertical, thirteen seconds, cut down from a longer take with the dead space removed, on-screen text naming each item, music underneath and captions placed clear of the existing graphics.

Thirteen seconds is honest about what the take contained. Vizard Agent kept the stretches that show the product and cut everything else, rather than padding towards a round number with footage that shows nothing.

When does this not work well?

This edit can only ever keep what the take actually contains, and a phone video of clothing filmed at home has a set of fairly predictable limits. Vizard Agent finds and keeps every usable second in the recording, and it cannot create the shot you did not get on the day.

How do you fix a result that came back wrong?

Name the moment or the text. Vizard Agent keeps the shot boundaries, the detail crops, the music analysis and the caption tests, so restoring a section or moving a label is a re-render rather than another pass through the take.

How does Vizard Agent compare to doing it yourself?

By hand this means trimming a phone video in a phone app, which handles the obvious dead space and not the half-second of fumbling between every shot. The captions then get placed by dragging them somewhere that looks free, and the sticker you added last week is exactly where they land.

By hand Vizard Agent
Finding dead space Trim the obvious gaps Shots split, every boundary checked
Knowing what is shown Remember filming it Garment details zoomed into and read
Caption placement Drag somewhere free Checked against the frame's existing graphics
The music Pick a trending sound Structure and rhythm analysed before choosing

Common questions

Does this work for other products? Yes. Anything filmed casually at home to show an object — homeware, tools, cosmetics, plants — has the same shape: a few usable seconds inside a lot of reaching and adjusting, and Vizard Agent finds them the same way.

How short will it end up? As short as the usable footage. This run delivered thirteen seconds from a longer take.

Will Vizard Agent name the items on screen? Yes, and it is worth telling it what they are rather than relying on the transcript.

Can Vizard Agent handle right-to-left captions? Yes. It sources a suitable typeface and tests the rendering on a frame first.

What about stickers I already added? Vizard Agent checks the frames and places the captions clear of them.

Why zoom into the garment at all? Because the seconds worth keeping are the ones where the product is actually visible, and that is not obvious from a thumbnail. Looking at the print and the fabric close up tells Vizard Agent which stretches are showing something and which are just movement.

Does Vizard Agent add music? Yes, and it analyses each candidate's rhythm before choosing rather than taking the first result.

Can I get a longer version? Only if the take supports it. Padding with footage that shows nothing defeats the format.

Does Vizard Agent check the finished video? Yes. It reviews every frame of the final version and re-rendered this one after fixing the caption placement.