How to turn several clips into a thank-you video
Upload your clips and photos to Vizard Agent and say who the video is thanking. Vizard Agent transcribes every clip at once, tiles all their frames into a single mosaic, and reads both before cutting — so it knows what was said and what was shown across the whole set rather than clip by clip.
What is the short version?
A thank-you montage is made from whatever people happened to film, which means the material is uneven, unsorted and mostly unwatchable in its raw state. The job is knowing what is in it. Everything after that is straightforward.
- Go to Vizard Agent and upload every clip and photo you have.
- Say who is being thanked and what for.
- Say the length, the shape and whether you want music or the original voices.
What do you need before you start?
Everything you collected, and the occasion. Vizard Agent handles unsorted material well — it inspects and transcribes all of it before deciding anything — so the useful thing you supply is context: who these people are and why the video exists at all.
- All the clips and photos. Any length, any orientation, any quality. Vizard Agent sorts out what it has.
- Who is being thanked. The organisation, the unit, the team, the person.
- What for. The specific thing. A montage without a reason reads as a collection of clips.
- Whether the voices matter. If people spoke on camera, say whether their words should carry the video or sit under music.
- The length and shape. A minute is usually right, and wide suits something shown at an event.
What do you type into Vizard Agent?
Say the occasion plainly. Vizard Agent works out what is in the footage on its own, and the thing it cannot see is the relationship the video is about — which is exactly what decides whether a clip belongs in the cut.
Prompt
Variants worth knowing:
- Music instead of voices. If the audio is poor or the clips are silent, say so and Vizard Agent scores it instead.
- Photos only. A montage works from stills alone if that is all you have.
- Subtitles. Worth asking for when the video will be shown at an event where the sound is unreliable.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real thank-you montage. Two of the steps run everything at once rather than one file at a time, which is why a large unsorted upload is not a large amount of waiting.
- Checks the format and details of every uploaded video.
- Extracts the technical information for each file.
- Runs automatic transcription on every clip in parallel rather than sequentially.
- Extracts frames from each video to review the visual content.
- Combines all those frames into a single mosaic for a fast overall review.
- Reads the mosaic to analyse what each video actually shows.
- Checks the working directory for anything else that arrived.
- Reads the transcription files to verify the words spoken in each clip.
Steps 3 and 5 are the pattern worth noticing. Vizard Agent parallelised the slow per-file work and then collapsed the results into one image and one set of transcripts, so the whole collection could be judged as a whole rather than remembered clip by clip.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1920x1080, H.264, 30fps, 59.2 seconds, AAC audio. Wide and a minute long, which suits something shown on a screen at an event rather than scrolled past on a phone.
Ask for a vertical version in the same conversation and Vizard Agent re-frames from the same analysis, following the people in each shot rather than cropping the centre.
The shape of the cut usually follows the material: a few longer spoken moments carrying the middle, shorter clips grouped around them, and the group shot at the end. Vizard Agent proposes that order from what the transcripts and the mosaic showed, and you can rearrange it by naming the moment rather than by editing anything.
How long it takes was not measured separately here. Comparable work runs a median of 28 to 38 minutes end to end, and the upper quartile is several times the median. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.
When does this not work well?
The footage was filmed by people who were not thinking about a montage at the time, and no amount of editing changes where the camera was pointed. Vizard Agent chooses well among what exists and cannot supply the shot nobody took, or the audio nobody could hear over the crowd.
- Uneven audio across clips. Phone clips from different rooms vary wildly in level. Vizard Agent balances them, and there is a point where music over the top is the honest answer.
- Vertical and wide mixed together. Some clips will be letterboxed or cropped whichever shape you choose. Say which you prefer.
- The reason has to come from you. Vizard Agent can see people smiling; it cannot know why that mattered.
- People in the footage did not consent to a montage. Filming at an event and publishing a video are different things, particularly where uniforms, children or patients are involved.
- Very short clips limit the pacing. A montage cut from three-second fragments feels frantic no matter how it is assembled.
How do you fix a result that came back wrong?
Name the clip or the moment. Vizard Agent keeps the transcripts, the frame mosaic and the assembled cut, so swapping a section works from an analysis it already built rather than re-inspecting every file, which is the expensive part when there are twenty of them.
- "Use the clip where he says thank you." Vizard Agent has the transcripts and can find it.
- "That person should not be in it." A re-cut excluding those clips.
- "Put the group shot at the end." A reorder, not a re-edit.
How does Vizard Agent compare to doing it yourself?
By hand this is opening twenty phone clips one at a time to remember which is which, then cutting. The remembering is the work, and it is why these videos usually get made the night before and look it.
| By hand | Vizard Agent | |
|---|---|---|
| Knowing what is in each clip | Watch them all | Transcribes and tiles them |
| Finding a spoken moment | Scrub each clip | Search the transcripts |
| Balancing the audio | Level each one by ear | Measured across the set |
| Reordering | Drag and re-render | Say the order |
Common questions
Do I need to sort the clips first? No. Vizard Agent inspects and transcribes all of them, so unsorted is the expected input.
Can it mix photos and video? Yes. Upload everything together and Vizard Agent sequences them as one set.
What if the audio is bad? Say so. Vizard Agent can balance it, or score the montage with music and keep only the strongest spoken moments.
Can it add subtitles? Yes, and it is worth it for anything shown at an event where the sound may not carry.
How many clips can it take? A lot. Vizard Agent transcribes them in parallel and reviews their frames as one mosaic, so twenty clips is not twenty times the wait of one.
Can I get both wide and vertical? Yes. Ask for both and the second comes from the same analysis rather than a fresh pass.