Vizard Agent

How to keep your title text out of the generated image

Last updated 2026-09-10 · 8 min read

Tell Vizard Agent that the titles are typography, not part of the picture. It builds each card as real type laid over the image rather than asking the generator to draw the words, so the spelling is yours — and then measures what each finished card actually runs for, because a card asset routinely comes back longer than the duration it was asked for.

What is the short version?

An image generator draws something that looks like writing. At a glance it reads as a title; at full size the letters are wrong, and they are wrong differently every time you regenerate. Type set over the image is simply correct.

  1. Go to Vizard Agent with your wording.
  2. Say the titles are type over the picture, not in it.
  3. Give the exact spelling you need.

What do you need before you start?

The wording and a sense of the look. The wording matters more than it sounds — a title that has to be right is exactly the thing a generated image gets wrong, and handing it over as text removes the problem rather than reducing it.

What do you type into Vizard Agent?

Tell Vizard Agent where the text actually lives. "Make me a title card" is ambiguous between a designed card and a generated picture of a card, and only one of those two things will spell your title correctly on every single render.

Prompt

Build the opening card and the chapter cards as real typography over the picture — do not put the words inside the generated image. The wording is exactly as I have written it. Three seconds each, and tell me what they actually run for once they are built.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real documentary explainer with an opening card, two chapter cards and a closing card. The last two steps are a timing fault that only shows up once everything is assembled.

  1. Checks the generation options and their prices before committing.
  2. Verifies the facts against institutional sources.
  3. Writes the narration first so the edit is paced to the real spoken duration.
  4. Generates the narration in one consistent voice and measures it.
  5. Generates the visual set and an instrumental bed together.
  6. Reviews the generated frames for consistency and artefacts.
  7. Generates word timings from the finished narration.
  8. Creates the opening card with exact typography, not embedded in imagery.
  9. Creates each chapter card the same way.
  10. Creates the closing source card.
  11. Builds the picture timeline from the stills and the cards.
  12. Finds the assembly is longer than intended because the cards carry their own durations.
  13. Measures each card's real duration rather than the one requested.
  14. Trims the cards to their intended holds and brings the timeline back to length.

Step eight is the whole article. Typography laid over the picture is under your control — the spelling, the face, the size, the position — and none of that survives being described to an image generator instead.

Steps twelve to fourteen are the trap next door. A card asset frequently renders at a different length from the one requested, and a timeline built by adding up the durations you asked for will not match the durations you got, so each card is measured after it exists.

What does the result look like?

Cards whose words are exactly the words you wrote, in a typeface you chose, at the size and position you set, holding for the time you asked — over generated imagery that is doing the job it is good at, which is everything except the letters.

Regenerate the picture and the title is still spelled correctly.

When does this not work well?

Text that has to sit inside the world of the shot is a different problem. A sign on a building, a label on a packet or writing on a screen is part of the image, and pulling it out into a type layer means compositing it back into perspective.

How do you fix a result that came back wrong?

Say which card and what about it. Vizard Agent keeps the card definitions and their measured durations, so a change of wording, size or hold is an edit to one card rather than a rebuild of the sequence.

How does Vizard Agent compare to doing it yourself?

Asking a generator for a title card is quicker and produces something that looks right in a thumbnail. The spelling is checked at full size, if at all, and it changes on every regeneration, so the check has to be repeated every time.

By hand Vizard Agent
The words Drawn by the generator Set as type
Spelling Different every render Exactly what you wrote
Typeface and size Whatever appears Chosen and consistent
Card durations Assumed as requested Measured after building
Regenerating the picture Risks the title Leaves the title alone

Common questions

Why can a generator not spell? It draws letter shapes rather than setting text; Vizard Agent sets text.

Can the card still be animated? Yes. Vizard Agent animates the type over the picture.

What about a background image? Generate it freely — Vizard Agent puts the words on top.

Will the wording be exact? Yes. Vizard Agent uses the text you supply, character for character.

Can I choose the typeface? Yes, and Vizard Agent checks it can render your script.

Why did my timeline come out long? Card assets carry their own durations; Vizard Agent measures them.

Can I set the hold precisely? Yes. Vizard Agent trims each card to the time you name.

What about text inside the scene? A different job. Vizard Agent has to composite that into perspective.

Can it match a card style I have? Yes. Send one and Vizard Agent measures it.

Will regenerating the image break the title? No. Vizard Agent keeps the title on a separate layer.

Can it do chapter cards too? Yes. Vizard Agent builds them all the same way.

What about the closing credits? Same approach, and Vizard Agent keeps the wording exact.

Does this cost more? Less, usually. Vizard Agent finds type cheaper than regenerating images.

Why not just regenerate until the spelling is right? Because it is right by accident when it happens, and the next render is a fresh roll of the dice.