Vizard Agent

How to make a faceless archaeological documentary short

Last updated 2026-08-21 · 6 min read

Tell Vizard Agent the subject and the tone and it builds the whole short. Vizard Agent records the narration first, extracts frame-accurate word timings, generates the excavation visuals to fit them, and mixes the voice, music and effects at measured levels rather than normalising each stem separately.

What is the short version?

A faceless documentary short is narration plus visuals plus a serious register. Everything hangs off the voice: the shots are cut to its word timings, the captions are aligned to it, and the music sits underneath at a level that was measured rather than guessed.

  1. Go to Vizard Agent and say what the short is about.
  2. Say the tone — serious archaeological science, not mystery-channel.
  3. Say the length and that it should be vertical and faceless.

What do you need before you start?

A subject and a register. Vizard Agent writes the narration, sources the voice and generates the visuals, so a paragraph is enough to start, and stating the register precisely is what keeps a science documentary from drifting into speculation.

What do you type into Vizard Agent?

Describe the register as carefully as the subject. Vizard Agent uses the tone to choose the voice, the music, the grade and the way the captions animate, so "serious archaeological science documentary" shapes far more of the result than it appears to.

Prompt

Create a [30] second vertical faceless documentary short about [subject]. Style: serious archaeological science documentary, ultra-realistic excavation environments. Narrated, with captions.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real documentary short. The sequence is the interesting part: the voice comes first and everything else is measured against it, which is the opposite of cutting picture and adding narration later.

  1. Checks what generation and editing tools are available, including effects, voices, music and captions.
  2. Searches for a narrator and generates the documentary narration.
  3. Transcribes its own narration to get precise word timings, then inspects the timestamps.
  4. Analyses the voiceover segments to plan the scene boundaries against them.
  5. Searches stock footage for excavation and laboratory scenes, then generates cinematic vertical visuals for the scenes it cannot source.
  6. Builds a preview grid of the generated scenes and inspects them before using any.
  7. Generates the music bed and the transition and scan sound effects.
  8. Builds the timed voiceover track, extracts frame-accurate word timings and generates animated captions aligned to them.
  9. Measures every audio stem in LUFS, mixes at exact un-normalised levels, renders the master, and reviews a tiled sheet of quality-control frames before delivery.

Step 9's mixing decision is worth stealing. Normalising each stem to the same target and adding them together buries the voice; measuring the stems and mixing to the balance you want keeps the narration on top where it belongs.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 30.0 seconds, AAC audio. Vertical, exactly thirty seconds, narrated throughout, with generated excavation visuals, animated captions aligned to the narration, documentary-style badges and overlays, and a closing card for the channel.

Landing on exactly thirty seconds happens because the narration was conformed to the timeline rather than the timeline stretched to the narration. Vizard Agent fits the read to the format when the format is the fixed thing.

When does this not work well?

A documentary makes claims about the real world, and generated visuals make no claims at all. Vizard Agent produces the documentary look convincingly, which raises the standard the script has to meet rather than lowering it, because the pictures will be believed.

How do you fix a result that came back wrong?

Say what is off — a fact, a shot or the pacing. Vizard Agent keeps the narration with its word timings, every generated scene, the overlays and the stem measurements, so a correction is a targeted re-render rather than a fresh build.

How does Vizard Agent compare to doing it yourself?

By hand this is writing a script, recording it, sourcing or generating a dozen separate shots, aligning the captions word by word, and then discovering that the music is sitting on top of the narration in the final mix.

By hand Vizard Agent
The order of work Picture first, voice later Voice first, everything timed to it
Caption alignment Nudge them by hand Frame-accurate word timings
The visuals Source or commission Sourced, then generated for the gaps
The mix Normalise and hope Stems measured, mixed to a chosen balance

Common questions

Are the visuals real footage? Generated, with stock footage used where it fits. They illustrate the subject rather than documenting a real site, and the script should be written knowing that.

Can it hit exactly thirty seconds? Yes. Vizard Agent conforms the narration to the timeline rather than the other way round.

Does it add captions? Yes, animated and aligned to frame-accurate word timings.

Why record the narration first? Because everything else is timed to it. Vizard Agent transcribes its own narration to get word timings and cuts the visuals and captions against them.

Can I supply the script? Yes. Give it and Vizard Agent records and builds to your words.

Does it source real footage too? Yes. Vizard Agent searches stock for excavation and laboratory scenes first and generates only what it cannot find.

Can it make a series? Yes. A second episode on a new subject follows the same build.

Will the narration sound like a documentary? Vizard Agent searches for a voice against the register you describe, so it is worth saying precisely.

Does it add overlays and badges? Yes. Vizard Agent builds the scan lines, data badges and closing card that this format uses.

How does it balance the mix? Vizard Agent measures every stem in LUFS and mixes to a chosen balance rather than normalising each one separately and hoping the voice survives.

Does it check the finished cut? Yes. Vizard Agent tiles quality-control frames from the render and reviews the layout, the captions and the composition before delivering anything.