Vizard Agent

How to add motion graphics to a video that already has captions

Last updated 2026-09-01 · 8 min read

Upload the video with its captions and logo already burnt in. Vizard Agent measures where the face, the logo and the caption band actually sit across the whole clip, maps the zones left empty around them, and composites transparent motion graphics into those zones at the right moments.

What is the short version?

Adding graphics to a video that already has captions is a layout problem before it is a design problem. There are only a few places on the frame that are genuinely free, and finding them requires measuring rather than guessing.

  1. Go to Vizard Agent and upload the finished video.
  2. Say which points in the talk deserve a graphic.
  3. Give your brand colours so the overlays look like they belong.

What do you need before you start?

The finished video and an idea of what needs illustrating. Vizard Agent finds the safe zones itself by measuring the existing layout, so what it needs from you is the editorial decision about which moments benefit from a graphic and which are better left alone.

What do you type into Vizard Agent?

Say what the video is and what should be emphasised. Vizard Agent works out the placement from the footage itself, so the useful instruction is editorial rather than positional — telling it where to put things is usually worse than letting it measure.

Prompt

Add motion graphic overlays on top of this pre-recorded video, near the speaker, to make it more engaging. The video already has [my logo pinned top-left] and [burnt-in word-level captions]. Topic: [account security, SIM swap attacks, carrier protection]. Brand styling: [dark navy] with [cyan] accents. Put graphics on the points worth reinforcing — do not cover the speaker, the logo or the captions.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real vertical talking-head video that already carried a logo and burnt-in captions. Steps eleven to thirteen are the whole method: measure where the face sits, mock up the free zones, then look at the mock-up.

  1. Probes the uploaded video and extracts frames to see the footage.
  2. Transcribes the speech and reads the full transcript.
  3. Reviews the frames and grabs a larger one to study the layout.
  4. Checks the motion-graphics tool and the available effect library.
  5. Looks at the client's logo file from a previous project.
  6. Checks the brand's colours so the overlays match.
  7. Measures where the face sits in the frame across the video.
  8. Mocks up where graphics can sit and reviews the placement zones.
  9. Fixes the safe zones as coordinates — beside the head, below the captions, right of the logo.
  10. Builds the beat map of which moment gets which graphic.
  11. Generates transparent overlays in the brand styling.
  12. Composites them onto the footage at the exact timestamps.

Step nine is what makes this reliable. The zones came out as actual proportions of the frame — the strip beside the head between roughly 13% and 40% down, the banner below the captions from 74% down, the bar to the right of the logo at the top — so every overlay lands somewhere provably empty.

What does the result look like?

The source this page is written from measured 2160x3840 at 24fps and ran 169.6 seconds — a vertical talking head with the client's logo pinned to the top-left corner and word-level captions already burnt in around two-thirds of the way down the frame, both of which the new graphics had to work around.

The graphics sit beside the speaker's head, in a banner below the caption band, and in the top strip to the right of the logo. Nothing overlaps anything. Because the overlays are transparent, the original footage plays through underneath rather than being covered by a card.

When does this not work well?

A busy frame has little room left in it, and graphics that fight the existing layout make a video worse rather than better. Vizard Agent tells you when there is nowhere clean to put something instead of dropping it on top of the captions.

How do you fix a result that came back wrong?

Say which overlay is wrong and what is wrong with it. Vizard Agent keeps the measured safe zones, the beat map of which moment gets which graphic, and every transparent asset, so a repositioned or retimed graphic re-composites without any of the others being regenerated.

How does Vizard Agent compare to doing it yourself?

By hand this means eyeballing the placement on a single frame, dropping the graphic there, and discovering at minute two that the speaker leans right into it. Vizard Agent measures where the face sits across the whole video before it decides where anything is allowed to go.

By hand Vizard Agent
Placement Judged from one frame Measured across the video
Safe zones Assumed Mapped as frame proportions
Overlay assets Opaque cards, usually Transparent, footage plays underneath
Timing Dragged on a timeline Placed against the transcript

Common questions

Do I need to remove my captions first? No. Vizard Agent measures around them and places the graphics elsewhere.

Will the graphics cover the speaker? No. Vizard Agent measures where the face sits before choosing any position.

Are they transparent? Yes. Vizard Agent generates overlays with alpha so your footage plays through.

Can it match my brand? Yes. Vizard Agent takes hex values and styles the overlays to them.

How many graphics should there be? Fewer than you think. Vizard Agent places them on the points worth reinforcing.

How does it know when to show them? From the transcript. Vizard Agent times each one to the line it illustrates.

What if the speaker moves around? It measures across the video. Vizard Agent picks zones that stay free throughout.

Can it use my logo? Yes. Vizard Agent picks up the existing logo and keeps the new graphics clear of it.

Does it work on landscape video? Yes. Vizard Agent maps the safe zones for whatever frame it is given.

Can I pick the positions myself? You can, but the measured zones are usually better. Vizard Agent will follow your positions if you insist.

Why not just rebuild the whole video? Because you already published this one, or it took a day to caption. Vizard Agent adds a layer to what exists rather than making you start again.

Can it add graphics to a long video? Yes. Vizard Agent maps the zones once and places graphics across the whole runtime.

Does it check the result? Yes. Vizard Agent reviews frames at each overlay to confirm nothing collides with the speaker, the logo or the captions.

What if my captions move around? Say so. Vizard Agent measures the caption band across the whole runtime rather than sampling one frame, and picks zones that stay clear throughout.

Can it also restyle the existing captions? That is a different job. Vizard Agent can strip and rebuild them, but adding graphics around them leaves your published look intact.