Vizard Agent

How to stop the punch-ins cycling too fast in a talking-head edit

Last updated 2026-09-05 · 8 min read

Tell Vizard Agent the reframing is cycling too quickly and that a punch-in may only start where a new sentence starts. It reads the sentence boundaries out of the transcript, plans far fewer framing changes, holds each one across a long passage, and keeps the talk as one unbroken audio track underneath.

What is the short version?

A talking head that changes framing every few seconds is exhausting to watch, and worse, the framing changes are usually cut in the audio as well, so the speaker's voice takes a small hit at every one of them. Both problems come apart the same way.

  1. Go to Vizard Agent with the edit that feels too busy.
  2. Say punch-ins may only begin at the start of a sentence.
  3. Say the audio must run continuously underneath.

What do you need before you start?

You need the talk itself and a clear statement of the symptom. You do not need to work out how many framing changes the video should have, because that number falls out of the sentence structure once the rule is applied — which is exactly why the rule is better than the number.

What do you type into Vizard Agent?

Ask for the rule rather than the count. "Fewer zooms" leaves Vizard Agent guessing at how many fewer, and a guess will be wrong in one direction or the other; "only at the start of a sentence" is a rule the transcript can actually be measured against.

Prompt

The framing changes too often and it is tiring to watch. Punch-ins must only start at the beginning of a new sentence, never mid-sentence, and each framing should hold for a long stretch. Keep the talk as one continuous audio track — do not cut the speech at the framing changes. Same total length.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real eight-minute talking head whose owner asked for exactly this. The listing step comes first, and everything after it is downstream of that list rather than of anybody's taste.

  1. Lists every sentence timing in the original transcript.
  2. Plans sparse framing changes anchored on those sentence starts.
  3. Plans slower punch-ins so each framing holds across a passage.
  4. Renders a forty-second test of the new framing only.
  5. Samples the test for punch-ins, captions and audio together.
  6. Checks the opening holds wide while the first graphic sits.
  7. Checks the first punch-in actually lands on a sentence start.
  8. Checks the test audio is continuous across that punch-in.
  9. Renders the full talk with the sparse punch-ins applied.
  10. Samples the recut at every graphic and punch-in beat.
  11. Checks each on-screen line against the words being spoken.
  12. Checks the full-length waveform for cuts that should not be there.
  13. Uploads the recut and watches it back end to end.

Step one is the part worth copying. The transcript already knows where the sentences are, so the framing plan can be derived from the speech rather than laid over the top of it on a timer — and a plan built that way never lands a zoom in the middle of a clause.

Step twelve is the proof. A framing change that was cut in the audio leaves a visible mark in the waveform, so checking the waveform of the finished piece is a direct test of the "continuous audio" instruction rather than a matter of listening hard and hoping.

What does the result look like?

A talk that still moves, with far fewer framing changes than before, each one arriving at a natural break in the speech and staying there long enough to stop registering as an effect. The voice underneath runs from beginning to end without a seam.

The best sign it worked is that nobody mentions the framing at all.

When does this not work well?

Sentence starts are only useful anchors when the speaker actually speaks in sentences, and plenty of good footage does not — a rambling take, a conversation, a heavily accented recording that the transcript struggles with, or a piece so short that holding a framing means never changing it.

How do you fix a result that came back wrong?

Point at the moment and say what is wrong with it, in ordinary words. Vizard Agent keeps the sentence list and the framing plan it derived from that list, so moving one punch-in or lengthening one hold is a change to the plan rather than a fresh pass over the whole video.

How does Vizard Agent compare to doing it yourself?

By hand, reframing on a long talk is usually done on feel, a zoom every so often to keep it alive, and the cut points end up wherever the editor happened to be scrubbing. It reads as restlessness, and the fix is normally described as "fewer zooms" rather than as a rule.

By hand Vizard Agent
Where framing changes Wherever it felt due At sentence starts, from the transcript
How long each holds Roughly even Across a whole passage
The audio at a change Often cut with the picture One continuous track
Proof it is continuous Listening The waveform of the final file
Fixing one moment Re-work the section Move one entry in the plan

Common questions

Why sentence starts specifically? Because a framing change there reads as punctuation. Vizard Agent uses the transcript to find them.

How many punch-ins will I get? As many as the rule allows. Vizard Agent does not work to a target count.

Can I ask for a minimum hold? Yes. Say the number of seconds and Vizard Agent plans around it.

Does the video get shorter? No, not unless you ask. Vizard Agent keeps the full runtime.

Will the captions still line up? Yes. Vizard Agent remaps the word timings onto the new cut.

Can graphics stay where they were? Yes. Vizard Agent times them to the spoken line, not to the framing.

What if I want one dramatic tight shot? Say where. Vizard Agent will hold it as long as you like.

Does it work on 4K footage? Yes. Punching in on a 4K source is why the framing stays sharp.

Can it undo an edit that is already too busy? Yes. Vizard Agent recuts from the original rather than from the busy version.

How does it know a punch-in cut the audio? It looks at the waveform of the render, where a cut is visible.

Can two speakers use this? Partly. Tell Vizard Agent to anchor on speaker changes as well.

What about vertical versions? Same rule. Vizard Agent applies it to the vertical crop too.

Will it tell me where the changes are? Yes. Ask Vizard Agent for the framing plan before rendering.

Why not just use fewer zooms? Because "fewer" is a guess and "at sentence starts" is a rule, and only one of the two produces the same result twice.