Vizard Agent

How to stop a narrator reading the timecodes out of your script

Last updated 2026-09-04 · 7 min read

Paste the script exactly as you wrote it and tell Vizard Agent not to say the timestamps. It treats the headers, timings and stage directions as instructions rather than lines, reads only the spoken words, and can deliver the narration and the music bed premixed as a single audio file.

What is the short version?

A working script is full of things nobody is supposed to say out loud: [0:00–0:40 — HOOK], [COLD OPEN], (pause). They are directions to the reader, and a narrator that reads them is technically doing exactly what you asked.

  1. Go to Vizard Agent and paste the script as it is.
  2. Say not to read the timestamps or the section headers.
  3. Say whether you want one premixed file or separate stems.

What do you need before you start?

The script and a voice. You do not need to strip your own formatting first — the structure in a script is useful information about pacing and emphasis, and removing it to protect against being read aloud throws that away.

What do you type into Vizard Agent?

Say it in the same breath as the delivery format. Both instructions are about what comes back rather than about the content, and putting them together means you do not discover the timestamp problem after listening to eleven minutes of narration.

Prompt

Create the voiceover and the background music as one audio file. Don't say the timestamps or the section headers. South African male voice, normal pace. Runtime target about 10 minutes.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real ten-minute narration built from a marked-up script. Step six is the small one that saves the job: it listens to a sample before committing to ten minutes of it.

  1. Checks the voiceover and music options.
  2. Searches for a voice matching the accent and gender you named.
  3. Searches the music library for the mood — several passes, refining.
  4. Checks the script's length against the runtime target.
  5. Formats and verifies the beats, separating directions from lines.
  6. Generates the first segment and analyses its tone and delivery.
  7. Analyses the delivery again before committing.
  8. Generates the remaining segments concurrently.
  9. Downloads the candidate music tracks and probes their durations.
  10. Checks each voiceover beat's duration against the plan.
  11. Analyses the music candidates for mood and style.
  12. Assembles the complete voiceover track.
  13. Converts everything to a uniform format and concatenates cleanly.
  14. Measures the loudness of the voice and the music separately.
  15. Builds a seamlessly crossfaded music bed for the full runtime.

Step five is where the timestamps stop being a problem. The markers are parsed as structure — this beat is the hook, this one runs forty seconds — so they shape the pacing instead of being spoken. That is a better outcome than deleting them, because the pacing information survives.

Steps six and seven are worth copying. Generating ten minutes of narration and then discovering the read is too theatrical is an expensive way to learn it; one segment, listened to, costs almost nothing.

What does the result look like?

One audio file with the narration sitting over a music bed, at the runtime you asked for, in the voice you named — and not a single spoken timecode or section header anywhere in it, because Vizard Agent treated all of those as directions rather than lines.

The music is crossfaded across the whole length rather than looped, which matters at ten minutes. A four-bar loop repeating for the duration of a long-form video is the other thing people notice immediately.

When does this not work well?

Some scripts are genuinely ambiguous about what is a line and what is a direction, and a bracket does not always mean "do not read this out". Where your own convention is unclear, it is worth telling Vizard Agent which is which rather than letting it decide.

How do you fix a result that came back wrong?

Say what you actually heard and roughly where. Vizard Agent keeps the parsed script, every segment's audio and the loudness measurements, so re-reading one line or changing the pace becomes a re-generation of that single segment rather than of the whole narration.

How does Vizard Agent compare to doing it yourself?

By hand the defence is to keep a second, stripped copy of every script — one to work from and one to read from. That works and it means every edit has to be made twice, which is where the two copies quietly stop matching.

By hand Vizard Agent
Protecting against read-aloud markers A stripped second copy Parsed as structure
Pacing information Lost in the stripped copy Used to time the beats
Checking the read After the full render On one sample segment
Music at length Looped Crossfaded across the runtime
Delivery Two files to mix One premixed file, or stems

Common questions

Do I need to clean the script first? No. Vizard Agent reads the markers as structure and leaves them unspoken.

What if my directions have no brackets? Say which lines they are. Vizard Agent asks when the convention is unclear.

Can I get the stems instead? Yes. Vizard Agent delivers voice and music separately if you prefer.

Will it hit my exact runtime? Close. Say if the number is hard and Vizard Agent adjusts the pace.

Can I hear a sample first? Yes — and you should. Vizard Agent generates one segment to check the read.

Can it do a specific accent? Yes. Name it and Vizard Agent auditions voices for it.

What about pauses written in? Honoured as directions. Vizard Agent times them rather than saying them.

Does the music loop? No. Vizard Agent crossfades a bed across the full length.

Can it read two characters? Yes, if you mark who says what.

Will it say brand names correctly? Tell Vizard Agent how they sound; it checks the delivery.

Can I use my own music? Yes. Send the track and Vizard Agent mixes to it.

How long can the script be? Long. Vizard Agent handles ten to eleven minutes routinely.

Can it split the file by section? Yes. Ask Vizard Agent for one file per beat if that suits your edit.

What if a segment sounds different? Say which. Vizard Agent re-reads that one against the others.

Does it check the levels? Yes. Vizard Agent measures the voice and the bed separately.

Why not just delete the markers? Because they carry your pacing — which beat is the hook, how long it runs — and a script stripped of them is a script that has lost its shape.