How to label what is on screen instead of captioning every word
Tell Vizard Agent you want labels rather than captions, and say what should be named. A recipe short does not need every word transcribed across the pan — it needs "gochugaru" to appear at the moment the gochugaru goes in, and nothing on screen the rest of the time.
What is the short version?
Captions are the default setting everywhere, and for a recipe they are the wrong one. Nobody reads a recipe off the subtitles; they watch what goes into the pot and want to know what it was, which is a job for labels rather than a transcript.
- Tell Vizard Agent you do not want running captions.
- Say what should be labelled — ingredients, quantities, steps.
- Say what the labels must never cover.
What do you need before you start?
The clips and the recipe. Vizard Agent reads the ingredients off the footage and the audio, but a list settles the spellings and the names you use, which matters when the same thing has three names in three languages.
- The clips. All of them, including the retakes.
- The ingredients. As you want them named.
- The length. Forty-five seconds, a minute.
- Where the labels must not sit. Your face, the pan, the logo.
- Any text you do want. An intro line, an ending.
What do you type into Vizard Agent?
Say "no main captions" explicitly. It is the kind of instruction that feels like it goes without saying and does not — a short with a voiceover will be captioned by default, and then the labels have to fight the subtitles for the same part of the frame.
Prompt
Variants worth knowing:
- "No main captions anywhere." The instruction people leave out.
- "As it goes in." Timing tied to the action, not the sentence.
- "Clear of my face." Where labels always end up otherwise.
What does Vizard Agent actually do?
Here is the Vizard Agent sequence on seven raw clips of a stew becoming a forty-five second short. The labelling work is inseparable from the edit, because a label is only right if it lands on the action it names.
- Reviews every clip — prep, cooking, ingredient additions, plating, tasting.
- Works out the cooking sequence and the ingredient names from the footage.
- Separates the usable recipe steps from the setup and retakes.
- Checks the seasoning sequence closely so each label matches what is actually being added.
- Plans the vertical crops around the cook and the pan, with bold ingredient text.
- Builds the fast-forward cooking sequence, checking each crop before rendering.
- Renders a vertical preview to check the food framing, the ingredient text and the voice-only sound.
- Matches the intro treatment to a previous video in the series, shifted to this dish's colour.
- Checks the full edit for cropped ingredients and overlapping text.
- Confirms the garlic label lands on the actual seasoning action.
- Tightens a shaky seasoning shot before the final render.
- Adjusts the labels so they stay clear of the cook's face, then renders the finished short.
Step ten is the check that distinguishes a labelled video from a decorated one. A label that appears two seconds after the ingredient went in is not just late, it names the wrong thing — the viewer reads "garlic" while watching the chilli powder go in.
Step four is upstream of that. Before anything can be labelled correctly, the order of the seasonings has to be established from the footage, which is not always the order the cook remembers.
Step twelve is the one every creator asks for on the second pass. Bold text large enough to read on a phone is also large enough to sit across your chin, so the positions are set against the frames rather than by a preset.
What does the result look like?
A short where the food fills the frame, each ingredient is named exactly as it goes in, the text is large enough to read without covering the cook, and there is no running transcript. The voice carries the instructions; the labels carry the nouns.
When does this not work well?
Some videos need the captions after all. Anything watched with the sound off, anything where the value is in what is said rather than what is shown, and anything an accessibility requirement covers should have proper subtitles — labels are an addition, not a replacement.
- Watched with sound off. Labels alone will not carry it.
- Accessibility requirements. Subtitles are not optional there.
- Talk-led videos. The words are the content.
- Many ingredients at once. Labels collide.
- Very fast cutting. A label needs a second on screen.
How do you fix a result that came back wrong?
Name the label and say what is wrong with it — its timing, its position or its wording. Vizard Agent keeps each label as a separate element rather than as part of one caption track, so it can move or retime a single one without disturbing the rest of the edit.
- "The garlic label is late." It is re-timed to the action it names.
- "It sits on my face." That label is repositioned for that shot.
- "Use the Korean name." The wording is changed everywhere it appears.
- "There are still captions." The subtitle track is removed and only labels kept.
How does Vizard Agent compare to doing it yourself?
By hand you either caption everything, because the tool does it in one click, or you place labels manually and spend the evening nudging them. The nudging is not the hard part; noticing that the third label is on the wrong shot is, because by then you have watched it fifteen times.
| By hand | Vizard Agent | |
|---|---|---|
| The default | Full captions | Labels, if you ask for them |
| Timing | Placed by eye | Matched to the action being named |
| Order | As remembered | Established from the footage |
| Position | Nudged | Set against the frames, clear of the face |
| Checking | On playback | Label by label, against the shot |
Common questions
Can I have both labels and captions? Yes, but they compete for the same space — Vizard Agent will place them apart.
What about quantities? Say so and Vizard Agent writes "2 tbsp gochugaru" rather than just the name.
Can the labels be in two languages? Yes. Vizard Agent can stack a second line under each one.
Will they match my other videos? Yes. Point Vizard Agent at a previous video and it matches the treatment.
Can it speed up the boring parts? Yes. Vizard Agent fast-forwards sections with the labels still landing correctly.
What if I misname an ingredient on camera? Tell Vizard Agent the correct name and the label will use it.
Can labels be handwritten-looking? Yes. Tell Vizard Agent the style you want.
Do labels work on a horizontal video? Yes, and Vizard Agent has more room to place them.
Can it label equipment too? Yes — pans, temperatures and times are all labelled the same way.
How long should a label stay up? Long enough to read; Vizard Agent clears it before the next arrives.
What if two ingredients go in together? Vizard Agent puts them on one label rather than stacking two.
Can I get a captioned version as well? Yes, from the same edit, so the two stay in step.
Will the voiceover still be heard? Yes. Labels are picture; Vizard Agent does not touch the audio.
Does it work for other kinds of video? Yes — assembly steps, tools, plants and products all label the same way.