Vizard Agent

How to build a thumbnail from a frame of your own video without the platform mark

Last updated 2026-09-15 · 8 min read

Ask Vizard Agent to take the thumbnail frame from the clean local master rather than from the version you uploaded. A frame grabbed from a published copy carries whatever the platform stamped on it, and that mark ends up permanently inside your thumbnail. It compares moments across the edit for expression and free space, pulls the chosen frame at full resolution, and checks the finished image at the size it will actually be seen.

What is the short version?

The best thumbnail is usually already in your video. The trap is where you take it from — a screenshot of the published version has a watermark on it, and once that is composited under your text it is not coming out.

  1. Go to Vizard Agent with the finished edit.
  2. Ask for the hero frame from the clean master.
  3. Ask to see the thumbnail at its real size.

What do you need before you start?

The local master and a sense of the channel it is going to. Vizard Agent can find a strong frame on its own, but a thumbnail that matches the rest of the channel gets chosen more often than one that is simply well made.

What do you type into Vizard Agent?

Tell Vizard Agent which copy to take the frame from. It is the one instruction here that cannot be repaired afterwards, because a platform watermark sitting under your composition means rebuilding the whole thumbnail from scratch rather than adjusting it.

Prompt

Make the thumbnail from a frame of this video — take it from the clean local master, not the uploaded version, so the platform's mark is not in it. Compare a few moments for expression and free space, then show me the result at the size it appears in a feed.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the sequence on a real channel thumbnail built from the video it belongs to. The frame choice is a comparison rather than a grab, and the source of the frame is settled before anything is composited.

  1. Compares several moments across the edit.
  2. Looks at the expressions and the free areas in each.
  3. Takes the chosen moment at high resolution.
  4. Checks the hero frame large before using it.
  5. Goes back to the clean local copy so no platform mark is included.
  6. Composes the thumbnail at the platform's own dimensions.
  7. Checks it at real size for legibility and face prominence.
  8. Looks at a recent thumbnail from the channel to match its palette.
  9. Rebuilds to that palette and removes anything off-brand.
  10. Sizes any coloured box to the text it contains.

Step five is the whole reason this article exists. Everything else can be adjusted afterwards; a watermark composited into the background cannot, and the mistake is easy to make because the uploaded copy is the one that is to hand.

Step two is what separates a good frame from a usable one. A brilliant expression with the face dead centre leaves nowhere for the text, so the choice is about where the empty space is as much as about the moment.

Step ten is a small thing that reads as amateur when it is wrong. A coloured rectangle that is wider or taller than the words inside it looks like a mistake, and matching the box to the text is the difference between a designed thumbnail and a decorated one.

What does the result look like?

A thumbnail that belongs to the video and to the channel, with a clean background, text that reads on a phone, and a face given room. Vizard Agent tells you which moment it came from, so you can ask for a different one by name.

When does this not work well?

Not every video contains its own thumbnail, and Vizard Agent will say when yours does not. A film with no faces in it, no strong single moment, or one whose best frame gives away the payoff all argue for building an image from scratch rather than grabbing one out of the edit.

How do you fix a result that came back wrong?

Say which part is wrong — the frame, the text or the layout around it. Because Vizard Agent built those as separate layers rather than as one image, replacing the text without disturbing the composition underneath is a single instruction rather than a rebuild.

How does Vizard Agent compare to doing it yourself?

By hand the frame gets screenshotted from wherever the video is open, which is usually the published page. The watermark is small enough to ignore in the editor and obvious once the thumbnail is live, and by then it is under three layers of text.

By hand Vizard Agent
The source frame A screenshot of the upload The clean local master
Choosing it The first good face Expression and free space
Resolution Whatever the screenshot gave Full resolution from the file
Checking Full size on a monitor At feed size
Matching the channel From memory Against a recent thumbnail

Common questions

Why not screenshot the published video? Because the platform's mark comes with it and ends up in your thumbnail.

Can it pick the moment itself? Yes. Vizard Agent compares several and shows you why it chose one.

What if the face is centred? Then there is nowhere for text. Vizard Agent weighs that in the choice.

Does it match my channel? Yes, if you point Vizard Agent at a recent thumbnail from it.

Will the text read on a phone? Vizard Agent checks at feed size, which is the only size that matters.

Can it change only the text later? Yes. Vizard Agent keeps the composition exactly as it was underneath.

What size does it build at? The platform's own dimensions. Tell Vizard Agent which platform it is for.

Can it remove something from the frame? Vizard Agent can, though finding a cleaner frame is usually easier.

What about a series? Vizard Agent can hold one layout and vary the frame per episode.

Does it need my logo? Only if your channel puts one on its thumbnails. Vizard Agent will match what is there.

Can it avoid spoilers? Yes. Tell Vizard Agent which moments are the payoff.

Will it upscale a soft frame? Vizard Agent can, but it will usually point out that a sharper moment exists.

How many options can I see? As many as you want; Vizard Agent builds them from the same master.

Why is the frame source so important? Because everything else is editable afterwards and a baked-in watermark is not.