Vizard Agent

How to turn the interviewer's question into on-screen text

Last updated 2026-09-22 · 8 min read

Tell Vizard Agent to mute the interviewer and put each question on screen as text. The format exists because a spoken question costs five seconds and a caption costs none — the reel keeps the expert's voice and loses the off-camera person nobody came to hear.

What is the short version?

Half of any recorded question-and-answer session is somebody off camera asking things. Written down on screen, those questions become the structure of the reel; left spoken, they are the single reason the reel runs too long for the platform.

  1. Tell Vizard Agent to mute the person asking.
  2. Say the questions should appear as animated captions.
  3. Say what the B-roll under the answers should show.

What do you need before you start?

The recording, and a sense of what the product or the subject is. Vizard Agent separates the questions from the answers in the transcript itself, but knowing what the B-roll is supposed to demonstrate is what decides which parts of your footage become inserts under the answers.

What do you type into Vizard Agent?

Describe the format rather than the steps. "Questions as text, answers as audio, B-roll under the answers" is a complete brief, and it is a format Vizard Agent can apply consistently across a series once you have agreed it.

Prompt

Edit this into a reel: mute the person asking the questions and put each question on screen as clear animated text instead. Keep the expert's answers audible and in focus. Put B-roll under the answers — the wound before, applying the product, and the result — and finish on a clear shot of the packaging.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the Vizard Agent sequence on a veterinary product question-and-answer session turned into a reel. Splitting the transcript by speaker comes first, because every later decision in the edit depends on knowing exactly who said what and when.

  1. Reviews the duration, the audio and the key frames.
  2. Checks the analysis tools before reading the conversation.
  3. Reads the dialogue and the pauses out of the video.
  4. Builds a visual log to find the wound footage and the demonstration.
  5. Reads the transcript to separate the questions from the expert's answers.
  6. Reads the visual log to fix the timecodes of the product and wound footage.
  7. Checks the styled caption format for the on-screen questions.
  8. Storyboards the opening and closing sections to find the clearest product close-ups.
  9. Finds the word-timing file so the question captions can be synchronised.
  10. Checks the packaging framing at the dosage explanation and the closing testimonial.
  11. Times six questions and builds the animated interview-style captions.
  12. Checks the audio channel format before reassembling the track with the expert's voice alone.

Step five is the whole job in one step. Once the transcript is split by speaker, the questions become text you can style and time, and the answers become the only audio the reel needs.

Step twelve is the part that goes wrong quietly. Muting one voice is not a volume change when both people are on the same track — the audio has to be rebuilt so the asker is absent rather than quiet, without leaving gaps where their words were.

Step ten is what makes it an advert rather than a clip. If the product is going to close the reel, the frame where its packaging reads clearly has to be found rather than hoped for.

What does the result look like?

A reel that moves question, answer, question, answer, with the questions readable on screen and the expert's voice the only one you hear. The demonstration plays under the answers rather than interrupting them, and the last shot shows the product clearly.

When does this not work well?

Some conversations do not survive losing a voice. If the interviewer is part of the dynamic, if the answers reference what the asker just said, or if both people talk over each other, muting one side leaves the other sounding like they are replying to nothing.

How do you fix a result that came back wrong?

Say which question or which insert is the problem. Because Vizard Agent builds the reel as a list of question-and-answer pairs rather than one timeline, it can restyle, retime or drop a single pair without rebuilding any of the others.

How does Vizard Agent compare to doing it yourself?

By hand this is fiddly rather than difficult: typing out six questions, timing six captions, ducking one voice and finding B-roll for each answer. It is an hour of work per reel, which is why most Q&A reels simply leave the interviewer's voice in.

By hand Vizard Agent
Questions Typed and timed by hand Split from the transcript and timed
The asker's voice Ducked, and still audible Rebuilt out of the track
B-roll Whatever is nearby Located from a visual log
Product shots Whichever frame you paused on Chosen where the packaging reads
A series Redone each time One format, applied consistently

Common questions

Can the interviewer stay visible? Yes, if you want — Vizard Agent mutes the voice without cutting the picture.

What if both voices are on one track? Vizard Agent separates them where it can, and will tell you when it cannot.

Can the questions be in another language? Yes. Vizard Agent translates the captions while the answers stay as spoken.

How many questions fit in a reel? Six is comfortable; Vizard Agent will suggest cutting if it runs long.

Can it match my caption style? Yes. Give Vizard Agent an example and it matches the treatment.

Does the expert need a microphone? It helps enormously — Vizard Agent has only that voice to work with.

Can it add my logo? Yes, and the closing product shot is the natural place for it.

What about a series of these? Vizard Agent applies the same format to each recording once it is agreed.

Can it pick the B-roll itself? Yes, from the visual log — say what each answer should show.

Will the answers be subtitled too? Yes if you want, though Vizard Agent often captions only the questions.

Can it keep one question spoken? Yes. Name the one to keep and Vizard Agent leaves that exchange intact.

What length should each answer be? Whatever the thought takes — Vizard Agent cuts on sentences, not seconds.

Does it work horizontally? Yes, and there is more room for the question text.

How do I check it? Read the captions with the sound off; they should make sense alone.