How to turn a podcast into short clips
Upload the episode to Vizard Agent and ask for a number of clips and a platform. Vizard Agent finds the moments that make sense without the surrounding hour, cuts each one to start and end cleanly, reframes it vertically around whoever is speaking, and returns them captioned and ready to post.
What is the short version?
Three moves, and the hard part — deciding which sixty seconds of an hour are worth anyone's attention — is the part Vizard Agent does. Reframing, captioning and topping and tailing each clip come along with it rather than being three more jobs afterwards.
- Go to Vizard Agent and upload the episode.
- Ask for a number of clips and name the platform.
- Keep what works, and ask for changes on the rest.
The rest of this page is what goes into each of those, and where the approach stops working.
What do you need before you start?
The recording, and a rough idea of how many clips you want. Everything else Vizard Agent works out. What matters most is the audio: clip selection is driven by what is said, so an episode where two people talk over each other on one shared microphone gives it far less to work with than one with clean separate tracks.
- Clear audio. Separate tracks per speaker are ideal; a single room mic with crosstalk is the hardest case.
- Video, if you have it. Audio-only works and comes back as captioned waveform clips, but talking-head video reframes far better vertically.
- A number. "The best 5 clips" gives Vizard Agent a bar to clear. "Some clips" does not.
- A destination. TikTok, Shorts and LinkedIn want different lengths and different opening seconds.
What do you type into Vizard Agent?
Vizard Agent ships this as a template, and its prompt is already the right shape: the source, a number, and a platform. The number matters more than it looks — it is the bar Vizard Agent has to clear, and asking for five forces a ranking that asking for "some" does not. Fill in the brackets and send it with the file attached:
Prompt
Three variants worth keeping:
When you know the topic that matters
When you want one clip, not a batch
When you have a multi-camera recording
What are the steps inside Vizard Agent?
Three, and there is no transcript-reading stage in the middle. Vizard Agent listens to the episode, picks the moments, cuts and reframes them and writes the captions as one request, so what comes back for review is a set of finished clips rather than a list of timecodes to go and edit.
- Upload the episode to the Vizard Agent chat with the prompt above.
- Pick the tier that matches the job. For clips going out under a brand, the product recommends Max; the free Flash tier is aimed at quick, simple edits.
- Reply about specific clips. "Clip 3 ends too early, give it another ten seconds" is a valid follow-up — you are re-cutting, not re-running.
What does the result look like?
A set of separate clips, each already cropped to the platform's aspect ratio, captioned, and topped and tailed so it starts on a sentence rather than mid-word. Vizard Agent keeps them all in the project, so a clip you rejected is still there if you change your mind, and asking for five more later does not mean re-uploading the episode.
It is not instant. Across podcast and interview projects actually cut this way, the median run takes about 38 minutes end to end and costs about 59 credits, with the middle half landing between 18 and 61 minutes and between 30 and 220 credits. A long episode sits at the upper end — the work scales with how much there is to listen to.
When does this not work well?
Clip-finding depends on the episode containing moments that survive being taken out of it — and no amount of processing creates one that is not there. Everything that disappoints here traces back either to that, or to audio Vizard Agent cannot cleanly follow, because the selection is driven by what is said. Six cases are worth knowing before you upload an hour of recording:
- No self-contained moments. A meandering two-hour conversation with no arguments, no stories and no strong opinions has nothing to extract. This is the most common disappointment, and no tool fixes it.
- Heavy crosstalk on one microphone. Overlapping speech on a shared track can be hard to attribute, and mis-attributed captions read as sloppy.
- Long setups. A brilliant point that needs ninety seconds of context does not fit a sixty-second clip. Ask for a longer format instead of forcing it.
- Strong accents plus poor audio together. Either alone is usually fine; together they can raise the caption error rate enough that you should proofread.
- Heavy jargon and names. Product names, drug names and people's names are what captions get wrong. Give a short list up front and accuracy improves.
- More than two or three speakers. Vertical reframing has to decide who is on screen, and a five-person panel can change speaker faster than a vertical crop follows comfortably.
If the episode genuinely has no standalone moments, the honest answer is that it is not a clips problem — it is a format problem, and better questions in the next recording fix it.
How do you fix a result that came back wrong?
Reply in the same Vizard Agent conversation instead of starting over. The episode and its transcript are still there, so a follow-up re-cuts rather than re-processing an hour of audio, and you are not charged again for the transcription you already paid for. Refer to clips by number and say what is wrong with the result rather than guessing at the cause.
| What you see | What to say next |
|---|---|
| A clip starts or ends mid-thought | "Clip 2 cuts him off — extend the end until he finishes the sentence" |
| The clips are all from one section | "Spread them across the whole episode, not just the first twenty minutes" |
| Captions get a name wrong | Give the spelling: "it is Kaley, not Kayleigh, throughout" |
| The crop cuts someone out of frame | "Keep both people in frame in clip 4 instead of following the speaker" |
| The moments are not the interesting ones | Name the subject: "prioritise anything about pricing" |
How does Vizard Agent compare to clipping by hand?
By hand this is a listening job before it is an editing job: play the episode, note timecodes, cut each clip, reframe it, caption it, then check the captions. The listening pass alone runs at real time — an hour of podcast costs an hour before any editing starts.
| Vizard Agent | Manual clipping | |
|---|---|---|
| Finding the moments | Part of the same request | A full listen at real time |
| Reframing vertical | Automatic, follows the speaker | Keyframed by hand per clip |
| Captions | Generated, editable | Typed or auto-generated then corrected |
| Time for 5 clips from a 1-hour episode | About 38 min (half fall in 18–61) | Half a day, realistically |
| Judgement about what is interesting | Directed by your instruction | Yours entirely |
Neither replaces the other. Vizard Agent wins on the mechanical work and on volume; a human editor still wins when the choice of moment is the whole point — a launch announcement, a sensitive interview.
Common questions
Does Vizard Agent need video, or is audio enough?
Audio alone works and comes back as captioned clips built around a waveform or a static frame. Video gives a much better result on social, because a face reframed vertically holds attention in a way a waveform does not.
How long should the clips be?
Name a length if you care; otherwise Vizard Agent picks one that suits the platform. As a rule the clip should be as long as the idea needs and not one sentence longer — a padded ninety seconds performs worse than a tight forty.
Can it caption in a different language than the audio?
Ask for it explicitly and say which language. Translation and dubbing are a separate request with their own constraints — there is a full guide to that linked below.
Will it use my brand fonts and colours for the captions?
Say what you want — "captions in white on black, bottom third, no emoji" — and Vizard Agent follows it. Without an instruction it picks a readable default rather than guessing at a brand it has not been shown.
Can I get more clips from the same episode later?
Yes. The episode stays in the project, so asking for five more is a new instruction rather than a new upload, and Vizard Agent avoids repeating the moments it already gave you.