Vizard Agent

How to add subtitles and a logo to a video

Last updated 2026-08-16 · 5 min read

Upload the video to Vizard Agent and ask for subtitles and your logo. Vizard Agent transcribes the speech with word-level timing, writes it with real punctuation rather than a stream of words, and burns the captions in at the style and position you asked for, with the logo placed where you put it.

What is the short version?

Automatic captions are everywhere and most of them are bad in the same two ways: no punctuation, and a style that fights the video they sit on. Both are worth stating in the prompt rather than accepting, because both are one clause of instruction.

  1. Go to Vizard Agent and upload the video.
  2. Say the caption style you want and that the punctuation should be correct.
  3. Say where the logo goes, and either upload it or name the brand.

What do you need before you start?

Audible speech and two decisions: how the captions should look, and where the logo sits. Vizard Agent looks at your actual frames before placing anything, so the more specific you are about what must stay visible, the less there is to correct afterwards.

What do you type into Vizard Agent?

Asking for correct punctuation by name is the highest-value clause in this prompt. It is the difference between subtitles and a word stream, and it is the thing people forget to ask for. Everything else here is placement, which Vizard Agent decides against your real frames rather than a fixed offset.

Prompt

Add elegant but clear subtitles with correct English grammar and punctuation, and put the [brand] logo in the top [right] corner.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real captions-and-logo job. The captioning half runs the way you would expect, off a word-level transcript. The logo half is the part people do not expect, because Vizard Agent goes and finds the asset rather than waiting for you to supply it.

  1. Inspects the video and looks at what else is in the workspace.
  2. Extracts sample frames across the video and views them as a contact sheet, so it knows what the picture looks like where captions will sit.
  3. Transcribes with word-level timing, then reads the full word-level data rather than a paragraph version of it.
  4. Searches the web for the brand's logo and fetches the company's own pages to find a clean copy.
  5. Places the captions and the logo against the frames it examined in step 2, rather than at a fixed offset.

Step 4 is worth knowing about. If the brand is public, you do not have to go and find the asset yourself before you start.

What does the result look like?

From the run this page is written from, probed on the delivered file: 2160x3840, H.264, 25fps, 88.12 seconds, audio preserved. That is a 4K vertical output at the source's 25fps, so Vizard Agent followed the material and the destination rather than a fixed preset.

If you need particular specs — a smaller file, a different frame rate, a subtitle file alongside the burned-in captions — say so in the prompt and Vizard Agent will deliver that instead.

It is not instant, and this category was not separately measured. Comparable work runs a median of 28 to 38 minutes end to end. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.

When does this not work well?

Captioning fails at two edges: unusual words, and frames with nowhere safe to put text. Vizard Agent handles both better when told about them in advance, and neither is something it can guess from the video alone — a surname and a product name look like ordinary speech until someone who knows spells them.

How do you fix a result that came back wrong?

Name the word or the moment. Vizard Agent keeps the word-level transcript, so fixing a spelling re-renders the captions from data it already has rather than transcribing the video again, which is both faster and cheaper than starting over.

How does Vizard Agent compare to doing it yourself?

The manual version is auto-caption in an editor, then read every line and fix the punctuation, then style them, then find and place the logo and check it does not collide with anything. The transcription is the free part; the proofreading and the placement are where the hour goes, and both scale with the length of the video.

By hand Vizard Agent
Transcription Auto, then proofread every line Word-level, punctuated on request
Style Set once, adjust per clip Described in words
Logo Find the file, place it, check every frame Searches for it, places against real frames
Fixing a name Edit each occurrence Say the spelling once

Common questions

Can I get a subtitle file instead of burned-in captions? Yes. Ask Vizard Agent for the file, or for both.

Will it caption in another language? Yes. Say which, and whether you want the original language as well.

Can it match the caption style of my previous videos? Yes, if they are in the same project or you upload one. Vizard Agent reads the typeface and placement off the earlier frames.

Can it caption a video with no speech? There is nothing to transcribe, but you can ask Vizard Agent for titles instead: give it the lines you want and where they should appear, and it sets them in the same style captions would have used.

What if a word is transcribed wrong? Give the correct spelling. Vizard Agent re-renders the captions from the same transcript rather than re-transcribing.