Vizard Agent

How to put pictures behind a long recording you cannot cut

Last updated 2026-09-22 · 8 min read

Tell Vizard Agent that the recording is fixed — not cut, not shortened, not re-voiced — and that the picture must run continuously underneath it. It plans the visuals from the transcript, then matches the picture track to the audio duration exactly, which is where this job actually gets difficult.

What is the short version?

Somebody recorded a talk and it has to go out as it is. The words are the point, the voice is the person's own, and the only thing you are adding is something to look at for twenty-seven minutes.

  1. Tell Vizard Agent the audio may not be altered in any way.
  2. Say there must be no black screen at any point.
  3. Ask for the cost before anything is generated.

What do you need before you start?

The recording and a view on the look. Vizard Agent reads the transcript to decide what each section should show, but saying whether you want generated scenes, stock footage, or plain typographic cards changes both the feel and the cost considerably.

What do you type into Vizard Agent?

Say "do not generate yet" if you want the estimate first. It is a reasonable thing to ask on a job this size, and it costs nothing to plan the work, count the scenes and price it before committing to any rendering at all.

Prompt

Before generating anything, confirm you can use the full 27-minute Arabic audio exactly as it is — no cutting, no shortening, no rewriting, no replacing the voice. Then a continuous 16:9 YouTube video with relevant visuals the whole way through and no black screen at any point. Tell me the estimated cost first and do not generate the video yet.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on a 27-minute recording given a full picture track underneath it. Choosing and making the visuals is the straightforward half of the job; the timing arithmetic that keeps the picture exactly as long as the speech is what takes the extra passes.

  1. Checks the file's duration and streams and samples the beginning, middle and end.
  2. Reviews the tools and their prices to build a realistic estimate before anything runs.
  3. Extracts a working transcript to understand the subject and how many scenes it needs.
  4. Summarises the transcript's sections to decide how dense the visuals should be.
  5. Prepares a scene plan and the timing file the captions will need.
  6. Runs a parallel batch of images and short motion clips to cover the whole episode.
  7. Checks what completed and what is missing before assembling anything.
  8. Searches for free stock of landscapes and skies to sit between generated scenes.
  9. Builds an Arabic caption file and fixes the font and the safe-area placement.
  10. Turns the stills and clips into a continuous 16:9 track timed to the audio length.
  11. Finds the picture running long, measures every clip's duration, and identifies which files caused it.
  12. Re-joins with a uniform time base, then with a frame-accurate filter that removes the timing gaps.
  13. Lays the original audio and the captions over the picture and confirms both stream durations match.
  14. Re-renders so the video ends with the audio, with no silent tail.

Steps eleven and twelve are the real work. Joining a few hundred clips end to end does not give you the sum of their durations — container rounding and mismatched time bases add fractions of a second each, and by the end the picture is minutes longer than the speech.

Step eight is a quiet cost decision. Generating every second of a 27-minute film is expensive and unnecessary; free stock between the generated scenes carries the reflective passages perfectly well.

Step two is why the estimate is trustworthy. Vizard Agent prices the actual scene count from the transcript rather than quoting a rate, and you see the number before anything is rendered.

What does the result look like?

One continuous video the exact length of the original recording: the voice untouched from first word to last, a picture at every moment, captions inside the safe area and readable, and no black frames or silent tail at the end.

When does this not work well?

Some recordings are not worth illustrating end to end. A rambling two-hour file, or one where the audio quality makes the speech hard to follow, produces a long video that nobody watches — and the honest advice is to cut it, which is exactly what you said you could not do.

How do you fix a result that came back wrong?

Say where and what you saw. Timing faults and picture faults are different repairs, and Vizard Agent fixes a duration mismatch in the join rather than by trimming the audio, which is the one thing it is not allowed to do.

How does Vizard Agent compare to doing it yourself?

By hand this is an afternoon of dragging stills onto a timeline, and a second afternoon spent discovering that the picture has ended up ninety seconds longer than the audio. The obvious fix is the one thing the brief forbids, which is shortening the speech to match.

By hand Vizard Agent
Planning By ear, as you go From the transcript's own sections
Visuals Whatever is to hand Generated and stock, mixed for cost
Duration Drifts clip by clip Measured, then re-joined frame-accurately
Black frames Found on the final watch Checked across the whole video
Cost Discovered afterwards Estimated before anything renders

Common questions

Will the voice be changed at all? No. Vizard Agent treats the audio as fixed and builds the picture to it.

Can it add captions in the same language? Yes, and at this length it is worth it — Vizard Agent places them in the safe area.

How much does 27 minutes cost? Ask first. Vizard Agent counts the scenes from the transcript and prices those.

Can it use my own photos? Yes. Supply them and Vizard Agent mixes them with anything it generates.

What about music underneath? Vizard Agent can add it, though under a talk it usually fights the voice.

Does every scene have to be generated? No. Vizard Agent mixes stock between generated scenes, which is cheaper.

Can it match the imagery to what is being said? Yes — that is what Vizard Agent's transcript pass is for.

What if the audio has long pauses? The picture keeps running; Vizard Agent does not cut to black.

Can it do a vertical version? Yes, though at this length a horizontal frame suits the material better.

Will it end exactly with the speech? Yes. Vizard Agent re-renders so there is no silent tail.

Can I approve the scene plan first? Yes. Ask Vizard Agent for it — on a long job that is the cheapest place to change your mind.

What if I want it shorter later? That is a different edit, and it does mean cutting the audio.

Does it work for a lecture series? Yes, and the look can be kept consistent across episodes.

How long does the render take? Long. The picture is built in pieces and joined, then checked end to end.