How to make a faceless archaeological documentary short
Tell Vizard Agent the subject and the tone and it builds the whole short. Vizard Agent records the narration first, extracts frame-accurate word timings, generates the excavation visuals to fit them, and mixes the voice, music and effects at measured levels rather than normalising each stem separately.
What is the short version?
A faceless documentary short is narration plus visuals plus a serious register. Everything hangs off the voice: the shots are cut to its word timings, the captions are aligned to it, and the music sits underneath at a level that was measured rather than guessed.
- Go to Vizard Agent and say what the short is about.
- Say the tone — serious archaeological science, not mystery-channel.
- Say the length and that it should be vertical and faceless.
What do you need before you start?
A subject and a register. Vizard Agent writes the narration, sources the voice and generates the visuals, so a paragraph is enough to start, and stating the register precisely is what keeps a science documentary from drifting into speculation.
- The subject. A site, a campaign, a period, a discovery.
- The tone. Serious documentary, and say it plainly.
- The length. Thirty seconds is the working size for this format.
- A channel name. For the end card and the badges.
- Any facts that must appear. Dates, place names, findings.
What do you type into Vizard Agent?
Describe the register as carefully as the subject. Vizard Agent uses the tone to choose the voice, the music, the grade and the way the captions animate, so "serious archaeological science documentary" shapes far more of the result than it appears to.
Prompt
Variants worth knowing:
- Data overlays. Scan lines and badges suit this format and Vizard Agent will build them.
- A call to action card. For a channel, at the end.
- A second episode. Ask and it is built the same way with a new subject.
What does Vizard Agent actually do?
Here is the order Vizard Agent worked in on a real documentary short. The sequence is the interesting part: the voice comes first and everything else is measured against it, which is the opposite of cutting picture and adding narration later.
- Checks what generation and editing tools are available, including effects, voices, music and captions.
- Searches for a narrator and generates the documentary narration.
- Transcribes its own narration to get precise word timings, then inspects the timestamps.
- Analyses the voiceover segments to plan the scene boundaries against them.
- Searches stock footage for excavation and laboratory scenes, then generates cinematic vertical visuals for the scenes it cannot source.
- Builds a preview grid of the generated scenes and inspects them before using any.
- Generates the music bed and the transition and scan sound effects.
- Builds the timed voiceover track, extracts frame-accurate word timings and generates animated captions aligned to them.
- Measures every audio stem in LUFS, mixes at exact un-normalised levels, renders the master, and reviews a tiled sheet of quality-control frames before delivery.
Step 9's mixing decision is worth stealing. Normalising each stem to the same target and adding them together buries the voice; measuring the stems and mixing to the balance you want keeps the narration on top where it belongs.
What does the result look like?
From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 30.0 seconds, AAC audio. Vertical, exactly thirty seconds, narrated throughout, with generated excavation visuals, animated captions aligned to the narration, documentary-style badges and overlays, and a closing card for the channel.
Landing on exactly thirty seconds happens because the narration was conformed to the timeline rather than the timeline stretched to the narration. Vizard Agent fits the read to the format when the format is the fixed thing.
When does this not work well?
A documentary makes claims about the real world, and generated visuals make no claims at all. Vizard Agent produces the documentary look convincingly, which raises the standard the script has to meet rather than lowering it, because the pictures will be believed.
- The visuals are generated, not archival. They illustrate; they do not document a real site.
- Facts need checking. Give Vizard Agent the dates and findings you want stated.
- Real sites and human remains deserve care. Some subjects are inappropriate to dramatise.
- The register can slip. Ask for serious science if you do not want mystery-channel narration.
- Thirty seconds is one idea. A complex excavation needs a series, not a longer short.
How do you fix a result that came back wrong?
Say what is off — a fact, a shot or the pacing. Vizard Agent keeps the narration with its word timings, every generated scene, the overlays and the stem measurements, so a correction is a targeted re-render rather than a fresh build.
- "That date is wrong." Corrected, re-recorded, and everything re-times to the new read.
- "The third shot looks wrong." Regenerated from the same scene plan.
- "The music is too loud under the voice." Re-mixed against the measured stems.
How does Vizard Agent compare to doing it yourself?
By hand this is writing a script, recording it, sourcing or generating a dozen separate shots, aligning the captions word by word, and then discovering that the music is sitting on top of the narration in the final mix.
| By hand | Vizard Agent | |
|---|---|---|
| The order of work | Picture first, voice later | Voice first, everything timed to it |
| Caption alignment | Nudge them by hand | Frame-accurate word timings |
| The visuals | Source or commission | Sourced, then generated for the gaps |
| The mix | Normalise and hope | Stems measured, mixed to a chosen balance |
Common questions
Are the visuals real footage? Generated, with stock footage used where it fits. They illustrate the subject rather than documenting a real site, and the script should be written knowing that.
Can it hit exactly thirty seconds? Yes. Vizard Agent conforms the narration to the timeline rather than the other way round.
Does it add captions? Yes, animated and aligned to frame-accurate word timings.
Why record the narration first? Because everything else is timed to it. Vizard Agent transcribes its own narration to get word timings and cuts the visuals and captions against them.
Can I supply the script? Yes. Give it and Vizard Agent records and builds to your words.
Does it source real footage too? Yes. Vizard Agent searches stock for excavation and laboratory scenes first and generates only what it cannot find.
Can it make a series? Yes. A second episode on a new subject follows the same build.
Will the narration sound like a documentary? Vizard Agent searches for a voice against the register you describe, so it is worth saying precisely.
Does it add overlays and badges? Yes. Vizard Agent builds the scan lines, data badges and closing card that this format uses.
How does it balance the mix? Vizard Agent measures every stem in LUFS and mixes to a chosen balance rather than normalising each one separately and hoping the voice survives.
Does it check the finished cut? Yes. Vizard Agent tiles quality-control frames from the render and reviews the layout, the captions and the composition before delivering anything.