How to cut the dead space out of a clothing showcase
Upload the take and tell Vizard Agent to cut everything that is not the product. Vizard Agent transcribes the speech, splits the take into shots, zooms into the garment details to see what is actually being shown, and rebuilds the video with the dead space removed and the captions placed clear of the frame's existing graphics.
What is the short version?
A product take recorded on a phone is mostly reaching, adjusting and finding the angle, with maybe eight usable seconds inside it. Cutting the dead space is the whole edit, and the rest is making those eight seconds look deliberate.
- Go to Vizard Agent and upload the take.
- Say to remove the dead space and keep what shows the product.
- Say the register and the language for the captions.
What do you need before you start?
The take and the product names. Vizard Agent finds the shot boundaries and the usable stretches itself, so nothing needs marking. Telling it what the items are is worth doing, because the on-screen text names them and a transcription of casual speech will not always get them right.
- The take. However rough, however long.
- The items. What each garment actually is.
- The register. Trendy and fast, or calm and considered.
- The language. It decides the typeface and the text direction.
- Any stickers already in the frame. So captions avoid them.
What do you type into Vizard Agent?
Say what to remove rather than what to keep. Vizard Agent reads "cut the extra parts" as an instruction to find every stretch where nothing is being shown, which is a more reliable brief than trying to list the moments you want.
Prompt
Variants worth knowing:
- Trendy music. It carries the pace when the take has no energy of its own.
- Text naming each item. Especially where the product is the point.
- A detail beat. A held shot of the print or the fabric.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real product showcase. The step people skip is the seventh: it zoomed into the shirt's print and the trousers' fabric to see what the video was actually showing before deciding which seconds to keep.
- Probes the video and extracts the speech transcript.
- Extracts frames and analyses the footage and the speech together.
- Reviews the audio track to see where the useful speech sits.
- Splits the take into shots and builds an overview of every cut.
- Reviews all the frames at once to see the whole take in one view.
- Searches for trendy music and analyses each candidate's structure and rhythm.
- Zooms into the garment details — the shirt's print, the trousers' fabric — to see what is being shown.
- Checks the first and last frames of every shot and reviews the sequence for continuity.
- Sources a typeface for the language and tests it on frame, cuts and joins the usable sections, then repositions the captions after finding them colliding with stickers already in the picture.
Step 9's ending is specific to this format. Product videos filmed for social often already have stickers or text burned in from the phone app, and captions added on top will land on them unless somebody checks the actual frames.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 12.97 seconds, AAC audio. Vertical, thirteen seconds, cut down from a longer take with the dead space removed, on-screen text naming each item, music underneath and captions placed clear of the existing graphics.
Thirteen seconds is honest about what the take contained. Vizard Agent kept the stretches that show the product and cut everything else, rather than padding towards a round number with footage that shows nothing.
When does this not work well?
This edit can only ever keep what the take actually contains, and a phone video of clothing filmed at home has a set of fairly predictable limits. Vizard Agent finds and keeps every usable second in the recording, and it cannot create the shot you did not get on the day.
- Colour on a phone is unreliable. A garment's real colour is hard to trust from casual footage.
- Detail needs a detail shot. If you never held on the fabric, there is no fabric beat to cut to.
- Existing stickers constrain the layout. Captions have to work around whatever is already burnt in.
- A short take stays short. Thirteen seconds of product is thirteen seconds.
- Right-to-left languages need font work. Vizard Agent sources and tests one; unusual scripts remain the risk.
How do you fix a result that came back wrong?
Name the moment or the text. Vizard Agent keeps the shot boundaries, the detail crops, the music analysis and the caption tests, so restoring a section or moving a label is a re-render rather than another pass through the take.
- "You cut the part showing the back." Restored from the mapped shots.
- "The caption sits on the sticker." Repositioned and re-checked, as here.
- "Hold on the print longer." Re-timed from the detail crop already extracted.
How does Vizard Agent compare to doing it yourself?
By hand this means trimming a phone video in a phone app, which handles the obvious dead space and not the half-second of fumbling between every shot. The captions then get placed by dragging them somewhere that looks free, and the sticker you added last week is exactly where they land.
| By hand | Vizard Agent | |
|---|---|---|
| Finding dead space | Trim the obvious gaps | Shots split, every boundary checked |
| Knowing what is shown | Remember filming it | Garment details zoomed into and read |
| Caption placement | Drag somewhere free | Checked against the frame's existing graphics |
| The music | Pick a trending sound | Structure and rhythm analysed before choosing |
Common questions
Does this work for other products? Yes. Anything filmed casually at home to show an object — homeware, tools, cosmetics, plants — has the same shape: a few usable seconds inside a lot of reaching and adjusting, and Vizard Agent finds them the same way.
How short will it end up? As short as the usable footage. This run delivered thirteen seconds from a longer take.
Will Vizard Agent name the items on screen? Yes, and it is worth telling it what they are rather than relying on the transcript.
Can Vizard Agent handle right-to-left captions? Yes. It sources a suitable typeface and tests the rendering on a frame first.
What about stickers I already added? Vizard Agent checks the frames and places the captions clear of them.
Why zoom into the garment at all? Because the seconds worth keeping are the ones where the product is actually visible, and that is not obvious from a thumbnail. Looking at the print and the fabric close up tells Vizard Agent which stretches are showing something and which are just movement.
Does Vizard Agent add music? Yes, and it analyses each candidate's rhythm before choosing rather than taking the first result.
Can I get a longer version? Only if the take supports it. Padding with footage that shows nothing defeats the format.
Does Vizard Agent check the finished video? Yes. It reviews every frame of the final version and re-rendered this one after fixing the caption placement.