How to show the thing you are describing instead of your face
Tell Vizard Agent to keep your recording as a small corner window and fill the rest of the frame with what you are talking about. Your audio stays as the spine, the transcript decides what each section shows, and the subtitles belong under the main picture rather than inside the little window.
What is the short version?
You recorded yourself explaining something. On its own that is a face talking to a camera; with the screen showing the thing you are describing and you in a corner, Vizard Agent turns the same recording into an explainer.
- Tell Vizard Agent to put your video in a corner as a small window.
- Say your original audio is kept and leads the edit.
- Say what should fill the rest — generated scenes, stock, or your own footage.
What do you need before you start?
Your recording, and a decision about what the rest of the frame is made of. Vizard Agent builds the background from your own transcript, so the more concrete your language is, the more specific the pictures behind you can be.
- Your recording. Any framing; it will be shrunk.
- Which corner. And how large the window is.
- What fills the frame. Generated, stock, or supplied.
- Music. Usually something quiet underneath.
- Where the subtitles sit. Under the main picture.
What do you type into Vizard Agent?
Say the window is small and that yours is the voice. Otherwise the request reads as a standard edit with B-roll, and what comes back is your video full-frame with cutaways rather than the layout you actually had in mind.
Prompt
Variants worth knowing:
- "As a small window." The layout, not a cutaway.
- "Keep my original audio." The spine of the edit.
- "Not inside the little window." Where captions must not go.
What does Vizard Agent actually do?
Here is the Vizard Agent sequence on a piece to camera rebuilt as a corner window over generated backgrounds. The transcript comes first in the order of work, because what you said decides everything that appears behind you.
- Checks the video's length and audio.
- Extracts representative frames to see the framing and mood.
- Confirms exactly what is said and when.
- Reviews the overall delivery and how the existing footage is composed.
- Checks the image, video and music generation options and their prices.
- Checks the background image and video specifications before generating anything.
- Works out how the generated images will be placed against the content.
- Generates the themed backgrounds and the piano bed in parallel.
- Reviews the generated images together for look and composition.
- Measures the original video's level and the music's level separately.
- Checks the music file's length and format.
- Assembles the background, the corner window and the audio into one video, then checks the frames and the waveform.
Step three is what makes the backgrounds relevant rather than decorative. Each section of speech gets a picture chosen for what it actually says, which is only possible once the words and their timings exist as text.
Step ten matters because you are now mixing three things: your voice, the music and whatever the generated footage brings with it. Measuring them separately is what keeps your voice in front of the other two.
Step five is a cost decision made before anything runs. Filling several minutes with generated backgrounds is the expensive part of this format, and it is worth seeing the number before committing to it.
What does the result look like?
Your recording in the corner at a size where you are still recognisable, the rest of the frame showing what you are describing, the subtitles readable under the main picture, and the music sitting under your voice rather than beside it.
When does this not work well?
Some talks should stay full-frame. If your expression is the content, if you demonstrate something with your hands, or if the subject is too abstract to picture, the corner window throws away the reason people are watching you in the first place.
- Expressive delivery. A small window loses the face.
- Demonstrations. Your hands are the point.
- Abstract subjects. Generic backgrounds read as wallpaper.
- Long videos. Generated backgrounds get repetitive.
- Poor lighting on you. The window makes it harder to see.
How do you fix a result that came back wrong?
Say whether the problem is the layout or the pictures behind it. Vizard Agent treats those as separate passes — the window's size and position is one job, what fills the rest of the frame is another — so a note about one leaves the other alone.
- "The window is too small." It is enlarged without moving the subtitles.
- "Those images repeat." New backgrounds are generated for the repeated sections.
- "The subtitles are in the window." They are moved under the main picture.
- "The picture jumps." The transitions between backgrounds are smoothed.
How does Vizard Agent compare to doing it yourself?
By hand the layout is the easy part and the backgrounds are the work: finding or making a relevant picture for every twenty seconds of talk, then timing each one to the sentence it belongs to. That is why most videos in this format reuse the same four stock clips.
| By hand | Vizard Agent | |
|---|---|---|
| Backgrounds | A handful, reused | Made for each section of the transcript |
| Timing | By feel | To the sentence they illustrate |
| Your audio | Re-levelled by ear | Measured against the music |
| Subtitles | Wherever the preset puts them | Under the main picture, out of the window |
| Cost | Discovered later | Checked before generating |
Common questions
How big should the window be? Vizard Agent sizes it so your expression still reads while the picture keeps the frame.
Can it be a circle? Yes. Tell Vizard Agent the shape and where it sits.
Will my audio stay untouched? Yes. Vizard Agent keeps your recording as the spine and mixes around it.
Can I supply the backgrounds? Yes. Vizard Agent places each one against the section it belongs to.
What if I move around in my recording? Vizard Agent can crop the window to keep you centred inside it.
Can the window move? Yes, though a fixed corner is easier to watch.
How long can this format run? A few minutes. Past that Vizard Agent will suggest varying the treatment.
Does it work horizontally? Yes, and Vizard Agent has more room for the illustration there.
Can it caption in another language? Yes, under the main picture as usual.
What about a logo? Vizard Agent can place one in the opposite corner to your window.
Will the images repeat? Not if you say so — Vizard Agent tracks what it has already used.
Can the music duck under my voice? Yes, and Vizard Agent measures both to set it.
Can I get a version without the window? Yes, the same edit with the background full-frame.
How do I check it? Watch it small. If you cannot tell it is you in the corner, the window is too small.