How to put pictures behind a long recording you cannot cut
Tell Vizard Agent that the recording is fixed — not cut, not shortened, not re-voiced — and that the picture must run continuously underneath it. It plans the visuals from the transcript, then matches the picture track to the audio duration exactly, which is where this job actually gets difficult.
What is the short version?
Somebody recorded a talk and it has to go out as it is. The words are the point, the voice is the person's own, and the only thing you are adding is something to look at for twenty-seven minutes.
- Tell Vizard Agent the audio may not be altered in any way.
- Say there must be no black screen at any point.
- Ask for the cost before anything is generated.
What do you need before you start?
The recording and a view on the look. Vizard Agent reads the transcript to decide what each section should show, but saying whether you want generated scenes, stock footage, or plain typographic cards changes both the feel and the cost considerably.
- The recording. Full length, as it is.
- What it is about. A talk, a sermon, a lecture.
- The look. Generated scenes, stock, cards, or a mix.
- The aspect ratio. 16:9 for YouTube, usually.
- Whether captions are wanted. Usually yes at this length.
What do you type into Vizard Agent?
Say "do not generate yet" if you want the estimate first. It is a reasonable thing to ask on a job this size, and it costs nothing to plan the work, count the scenes and price it before committing to any rendering at all.
Prompt
Variants worth knowing:
- "Exactly as it is." The constraint that defines the job.
- "No black screen at any point." What "continuous" actually means.
- "Do not generate yet." Plan and price first.
What does Vizard Agent actually do?
Here is the Vizard Agent sequence on a 27-minute recording given a full picture track underneath it. Choosing and making the visuals is the straightforward half of the job; the timing arithmetic that keeps the picture exactly as long as the speech is what takes the extra passes.
- Checks the file's duration and streams and samples the beginning, middle and end.
- Reviews the tools and their prices to build a realistic estimate before anything runs.
- Extracts a working transcript to understand the subject and how many scenes it needs.
- Summarises the transcript's sections to decide how dense the visuals should be.
- Prepares a scene plan and the timing file the captions will need.
- Runs a parallel batch of images and short motion clips to cover the whole episode.
- Checks what completed and what is missing before assembling anything.
- Searches for free stock of landscapes and skies to sit between generated scenes.
- Builds an Arabic caption file and fixes the font and the safe-area placement.
- Turns the stills and clips into a continuous 16:9 track timed to the audio length.
- Finds the picture running long, measures every clip's duration, and identifies which files caused it.
- Re-joins with a uniform time base, then with a frame-accurate filter that removes the timing gaps.
- Lays the original audio and the captions over the picture and confirms both stream durations match.
- Re-renders so the video ends with the audio, with no silent tail.
Steps eleven and twelve are the real work. Joining a few hundred clips end to end does not give you the sum of their durations — container rounding and mismatched time bases add fractions of a second each, and by the end the picture is minutes longer than the speech.
Step eight is a quiet cost decision. Generating every second of a 27-minute film is expensive and unnecessary; free stock between the generated scenes carries the reflective passages perfectly well.
Step two is why the estimate is trustworthy. Vizard Agent prices the actual scene count from the transcript rather than quoting a rate, and you see the number before anything is rendered.
What does the result look like?
One continuous video the exact length of the original recording: the voice untouched from first word to last, a picture at every moment, captions inside the safe area and readable, and no black frames or silent tail at the end.
When does this not work well?
Some recordings are not worth illustrating end to end. A rambling two-hour file, or one where the audio quality makes the speech hard to follow, produces a long video that nobody watches — and the honest advice is to cut it, which is exactly what you said you could not do.
- Very long recordings. Cost and attention both scale badly.
- Poor audio. Pictures cannot rescue speech you cannot hear.
- Abstract subjects. Generic imagery starts to feel like wallpaper.
- Material with strict imagery rules. Say so up front.
- Recordings that would be better cut. The constraint is yours to set.
How do you fix a result that came back wrong?
Say where and what you saw. Timing faults and picture faults are different repairs, and Vizard Agent fixes a duration mismatch in the join rather than by trimming the audio, which is the one thing it is not allowed to do.
- "There is a black frame at 14:02." That join is rebuilt with a frame-accurate filter.
- "The video runs longer than the audio." The picture track is re-timed to the audio length.
- "The captions are outside the frame." They are replaced inside the safe area.
- "This section's imagery is wrong." Those scenes are regenerated from that part of the transcript.
How does Vizard Agent compare to doing it yourself?
By hand this is an afternoon of dragging stills onto a timeline, and a second afternoon spent discovering that the picture has ended up ninety seconds longer than the audio. The obvious fix is the one thing the brief forbids, which is shortening the speech to match.
| By hand | Vizard Agent | |
|---|---|---|
| Planning | By ear, as you go | From the transcript's own sections |
| Visuals | Whatever is to hand | Generated and stock, mixed for cost |
| Duration | Drifts clip by clip | Measured, then re-joined frame-accurately |
| Black frames | Found on the final watch | Checked across the whole video |
| Cost | Discovered afterwards | Estimated before anything renders |
Common questions
Will the voice be changed at all? No. Vizard Agent treats the audio as fixed and builds the picture to it.
Can it add captions in the same language? Yes, and at this length it is worth it — Vizard Agent places them in the safe area.
How much does 27 minutes cost? Ask first. Vizard Agent counts the scenes from the transcript and prices those.
Can it use my own photos? Yes. Supply them and Vizard Agent mixes them with anything it generates.
What about music underneath? Vizard Agent can add it, though under a talk it usually fights the voice.
Does every scene have to be generated? No. Vizard Agent mixes stock between generated scenes, which is cheaper.
Can it match the imagery to what is being said? Yes — that is what Vizard Agent's transcript pass is for.
What if the audio has long pauses? The picture keeps running; Vizard Agent does not cut to black.
Can it do a vertical version? Yes, though at this length a horizontal frame suits the material better.
Will it end exactly with the speech? Yes. Vizard Agent re-renders so there is no silent tail.
Can I approve the scene plan first? Yes. Ask Vizard Agent for it — on a long job that is the cheapest place to change your mind.
What if I want it shorter later? That is a different edit, and it does mean cutting the audio.
Does it work for a lecture series? Yes, and the look can be kept consistent across episodes.
How long does the render take? Long. The picture is built in pieces and joined, then checked end to end.