How to add subtitles and a logo to a video
Upload the video to Vizard Agent and ask for subtitles and your logo. Vizard Agent transcribes the speech with word-level timing, writes it with real punctuation rather than a stream of words, and burns the captions in at the style and position you asked for, with the logo placed where you put it.
What is the short version?
Automatic captions are everywhere and most of them are bad in the same two ways: no punctuation, and a style that fights the video they sit on. Both are worth stating in the prompt rather than accepting, because both are one clause of instruction.
- Go to Vizard Agent and upload the video.
- Say the caption style you want and that the punctuation should be correct.
- Say where the logo goes, and either upload it or name the brand.
What do you need before you start?
Audible speech and two decisions: how the captions should look, and where the logo sits. Vizard Agent looks at your actual frames before placing anything, so the more specific you are about what must stay visible, the less there is to correct afterwards.
- Audible speech. Everything downstream is the transcript Vizard Agent makes.
- A style in words. "Elegant and clear", "big and bold, two words at a time", "small, bottom, unobtrusive". These produce genuinely different videos.
- The logo, or the brand name. Upload the file if you have it. If you do not, name the company and Vizard Agent can search for it.
- A position and a size. Top left, bottom right, and roughly how large.
- Anything that must not be covered. A face, a chart, an existing lower third. Say it, and Vizard Agent will work around it.
What do you type into Vizard Agent?
Asking for correct punctuation by name is the highest-value clause in this prompt. It is the difference between subtitles and a word stream, and it is the thing people forget to ask for. Everything else here is placement, which Vizard Agent decides against your real frames rather than a fixed offset.
Prompt
Variants worth knowing:
- Word-by-word pop. "Captions one or two words at a time, centred, bold." The style short-form feeds expect, and a different job from subtitles.
- Subtitles, not captions. "Bottom, small, unobtrusive, do not cover the speaker."
- Another language. "Captions in Spanish." That is a translation as well as a transcription, so tell Vizard Agent whether you want both languages or only one.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real captions-and-logo job. The captioning half runs the way you would expect, off a word-level transcript. The logo half is the part people do not expect, because Vizard Agent goes and finds the asset rather than waiting for you to supply it.
- Inspects the video and looks at what else is in the workspace.
- Extracts sample frames across the video and views them as a contact sheet, so it knows what the picture looks like where captions will sit.
- Transcribes with word-level timing, then reads the full word-level data rather than a paragraph version of it.
- Searches the web for the brand's logo and fetches the company's own pages to find a clean copy.
- Places the captions and the logo against the frames it examined in step 2, rather than at a fixed offset.
Step 4 is worth knowing about. If the brand is public, you do not have to go and find the asset yourself before you start.
What does the result look like?
From the run this page is written from, probed on the delivered file: 2160x3840, H.264, 25fps, 88.12 seconds, audio preserved. That is a 4K vertical output at the source's 25fps, so Vizard Agent followed the material and the destination rather than a fixed preset.
If you need particular specs — a smaller file, a different frame rate, a subtitle file alongside the burned-in captions — say so in the prompt and Vizard Agent will deliver that instead.
It is not instant, and this category was not separately measured. Comparable work runs a median of 28 to 38 minutes end to end. Across all projects the median cost by tier is Flash 47, Pro 55, Max 242, Ultra 263 credits.
When does this not work well?
Captioning fails at two edges: unusual words, and frames with nowhere safe to put text. Vizard Agent handles both better when told about them in advance, and neither is something it can guess from the video alone — a surname and a product name look like ordinary speech until someone who knows spells them.
- Names, jargon and product spellings need help. A transcript will guess at a surname or a drug name. If a word matters, write it in the prompt and Vizard Agent will use your spelling.
- Busy frames leave nowhere safe. Sometimes there is no position where two lines of text do not cover something. Say what must stay visible and accept smaller type.
- A logo found online may not be your current one. If your brand changed recently, upload the file rather than letting Vizard Agent search for it.
- Burned in is burned in. Captions rendered into the picture cannot be switched off later. If you need them optional, ask Vizard Agent for a subtitle file as well.
- Timing is only as good as the audio. Word-level timing comes from the speech itself, so a clip with music mixed loudly over the voice will drift. Cleaning the audio first, or asking Vizard Agent to do it in the same run, fixes the captions as a side effect.
- Heavy accents and overlapping speech. Two people talking at once produces one transcript. Expect to correct names and crosstalk.
How do you fix a result that came back wrong?
Name the word or the moment. Vizard Agent keeps the word-level transcript, so fixing a spelling re-renders the captions from data it already has rather than transcribing the video again, which is both faster and cheaper than starting over.
- "You spelled the surname wrong." Give the spelling; Vizard Agent re-renders from the same transcript.
- "The captions cover the chart at 0:30." It moves them for that stretch.
- "Too big." Style adjustments do not touch the transcript at all.
How does Vizard Agent compare to doing it yourself?
The manual version is auto-caption in an editor, then read every line and fix the punctuation, then style them, then find and place the logo and check it does not collide with anything. The transcription is the free part; the proofreading and the placement are where the hour goes, and both scale with the length of the video.
| By hand | Vizard Agent | |
|---|---|---|
| Transcription | Auto, then proofread every line | Word-level, punctuated on request |
| Style | Set once, adjust per clip | Described in words |
| Logo | Find the file, place it, check every frame | Searches for it, places against real frames |
| Fixing a name | Edit each occurrence | Say the spelling once |
Common questions
Can I get a subtitle file instead of burned-in captions? Yes. Ask Vizard Agent for the file, or for both.
Will it caption in another language? Yes. Say which, and whether you want the original language as well.
Can it match the caption style of my previous videos? Yes, if they are in the same project or you upload one. Vizard Agent reads the typeface and placement off the earlier frames.
Can it caption a video with no speech? There is nothing to transcribe, but you can ask Vizard Agent for titles instead: give it the lines you want and where they should appear, and it sets them in the same style captions would have used.
What if a word is transcribed wrong? Give the correct spelling. Vizard Agent re-renders the captions from the same transcript rather than re-transcribing.