How to put music under a long interview only where it helps
Tell Vizard Agent the cut is locked and only the audio changes, then ask for music where the story earns it and silence where it does not. It scores from the transcript it already has, keeps the music far under the voice, and leaves the picture stream untouched so nothing else can shift.
What is the short version?
A long spoken piece is the hardest thing to score. Wall-to-wall music flattens it, two isolated cues in an hour do nothing, and either mistake is easy to make from a one-line brief. Vizard Agent places cues by what is being said, then shows you where they landed.
- Tell Vizard Agent the edit is approved and this is audio only.
- Ask for music where the account carries tension, and silence elsewhere.
- Check the cues from the review audio rather than watching the whole thing.
What do you need before you start?
The finished edit and a transcript of it. If Vizard Agent already cut the piece it has both, and saying so is what keeps it from analysing the footage again — on an hour of material that re-analysis is the bulk of the work and none of it changes the result.
Also decide how present the music should be. "Underneath, never noticed" and "carries the scene" are different jobs, and the second one is rarely what an interview wants.
What do you type into Vizard Agent?
Lock the picture in the first sentence, then describe where music belongs in terms of the content rather than the clock. A brief that asks for "background music" over an hour of speech gets answered as a bed, which is almost never what the piece needs to work properly.
The 66-minute cut is approved — audio only, do not rebuild the edit or re-analyse the picture. Reuse the existing transcript. Add restrained cinematic music where the account reaches tension or reflection, across the whole runtime, with real silence in between. Keep it far under the voice and duck it whenever he speaks.
The phrase doing the most work there is "across the whole runtime, with real silence in between". It rules out both failure modes at once.
What does Vizard Agent actually do?
It reads the transcript for where the story turns rather than laying music on a schedule. That is why the transcript matters more than the footage here: the cue points are in what is said, and the picture has nothing to tell you about them.
In the session behind this article it:
- Reused the existing finished timeline and transcript instead of reprocessing the source.
- Placed cues at developing tension, at the high-stakes passages, and at the reflective ones — with gaps where the words should stand alone.
- Held the music at roughly 15–20% of the perceived dialogue level and ducked it automatically under speech.
- Faded every cue in and out, so no cue starts or stops abruptly.
- Encoded the new mix and muxed it against the existing picture, which was never re-encoded.
It also built two review files: an overview plotting the dialogue at unchanged gain against the added music on the same scale, with bands marking the scored passages, and a listening reel holding each cue's entrance, middle and fade-out back to back.
What does the result look like?
The same film you approved with a score underneath part of it. The picture is bit-for-bit what it was, the dialogue is at its original level, and the music appears and leaves in a way you have to look for to notice.
The overview picture is the part worth keeping. It shows at a glance which stretches got music and which were left alone, and the unchanged dialogue trace is the proof that Vizard Agent did not touch anything else while it was in there.
When does this not work well?
When the piece is testimony rather than storytelling. Music under a first-hand account can read as the edit telling you how to feel about it, and on serious subject matter that is a real editorial choice rather than a finishing touch — worth making deliberately, not by default.
It also goes wrong when "restrained" is the only instruction. In that session the first pass put a handful of cues across an hour because restraint was the brief, and the correction was explicit: do not limit the number, score wherever the story supports it, just do not run continuously.
And a piece with inconsistent source audio fights you. If the voice level already wanders, Vizard Agent is ducking against a moving target and the music will feel uneven even when it is not.
How do you fix a result that came back wrong?
Say whether there is too little or too much, because those are opposite corrections and the words for them overlap. "More cues across the runtime, still with gaps" moves Vizard Agent in one direction; "fewer, only the strongest moments" moves it the other, and neither is a volume change.
If a single cue is wrong, name it from the review audio. Each excerpt maps back to a timecode in the master, so you can ask for one entrance moved or one fade lengthened without reopening the rest of the mix.
If the music is audible rather than felt, ask for a lower target as a percentage of dialogue level. Vizard Agent treats that as a number rather than a mood, which is why it is the reliable thing to correct.
How does Vizard Agent compare to doing it yourself?
Scoring an hour by hand means sitting through it repeatedly, and most of that time is spent confirming the passages you already decided to leave alone. Working from the transcript puts the decision where the information is and skips the rest.
The habit worth copying is the listening reel. Reviewing every cue's entrance and fade back to back takes minutes and catches the abrupt ones immediately, while watching the full hour catches them only if your attention happens to be in the right place.
Common questions
Will the picture be re-rendered? No. Vizard Agent muxes the new audio against the existing video stream.
Does it need to watch the footage again? No. Vizard Agent works from the transcript it already has.
How loud should the music be? Roughly 15–20% of the perceived dialogue level, or quieter.
Does it duck under speech? Yes. Vizard Agent ducks automatically, so the voice stays dominant.
Will cues start abruptly? No. Vizard Agent fades every cue in and out.
How many cues in an hour? As many as the story supports. Tell Vizard Agent "across the runtime".
Can I see where the music is? Yes. Vizard Agent plots the scored passages against the dialogue.
Do I have to watch the whole thing to check it? No. The listening reel holds each cue's entrance, middle and fade.
Can I move one cue? Yes. Name it from the review audio and Vizard Agent adjusts that one.
Will the dialogue level change? No. Vizard Agent leaves the original audio at its own gain.
Should an interview have music at all? Sometimes not. Vizard Agent will leave passages bare if you ask it to.
Can I ask for silence in specific sections? Yes, and it is worth telling Vizard Agent explicitly.
Does this cost less than a rebuild? Yes. Vizard Agent skips the analysis and the picture render.
What if the voice level wanders? Fix that first. Ducking against uneven dialogue never settles.