Vizard Agent

How to reframe a screen recording to follow the action rather than the cursor

Last updated 2026-09-10 · 8 min read

Tell Vizard Agent to follow what the user is doing rather than where the pointer is. It crops each section of the recording to the part of the screen where that step actually happens, changes the crop as the tutorial moves down the screen, and checks the beginning of each gesture — the tap that starts a drag — is still inside the frame.

What is the short version?

Centring on the cursor is the obvious way to reframe a screen recording and it produces a video that swims. The pointer moves constantly and the action does not; following the action means a small number of deliberate crops, each held while a step happens.

  1. Go to Vizard Agent with the recording.
  2. Say the crop should follow the action, not the cursor.
  3. Give the sections and which part of the screen each needs.

What do you need before you start?

The recording and a rough breakdown. You do not need exact frames — "the first part is the tap at the top, then the middle of the screen, then the bottom half" is enough for a first pass, and the timings get refined against what is actually there.

What do you type into Vizard Agent?

Say what the viewer needs to see. Asking to keep the cursor centred is a precise instruction and the wrong one — it describes the mechanism rather than the goal, and the mechanism is what makes the result unwatchable.

Prompt

Crop this to a square. Do not follow the white pointer dot — follow what the user is doing. From the start to about six seconds, keep the tap at the top. From ten to fifteen, use the lower part of the screen. From twenty-two to the end, the upper part. Make sure the tap that starts each drag is still in frame.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real phone screen recording that had to become a square tutorial. The crop changes several times, and each change is tied to a step rather than to a movement of the pointer.

  1. Checks the recording's shape and the space each step occupies.
  2. Builds a first version centred on the pointer, as asked.
  3. Reviews it and finds the framing swims with the cursor.
  4. Rebuilds the crop to follow the action instead.
  5. Sets a crop per section — top, lower half, upper half.
  6. Checks the beginning of each section for a missing tap.
  7. Extends the first section where the opening tap was cropped away.
  8. Checks a transition that stutters between two crops.
  9. Smooths the move between one crop and the next.
  10. Re-checks each section against what the user is doing there.
  11. Renders the final shape and reviews it end to end.

Step six is the check that saves the tutorial. A crop chosen for where a drag ends will frequently miss where it started, and a viewer who does not see the finger land has no idea what began the gesture they are now watching.

Step two and three are worth leaving in the account. The pointer-centred version gets built exactly as asked and then rejected on sight, which is often the fastest way to establish that the obvious approach is the wrong one — it costs one render and settles the question for good.

Step nine matters because the crops move. Jumping from one part of the screen to another between sections reads as a glitch; easing between them reads as a camera move, and it costs a few frames at each change.

What does the result look like?

A tutorial in the shape you asked for, where each step fills the frame while it is actually happening, the crop Vizard Agent chose moves deliberately between one step and the next, and every gesture is visible from the moment the finger lands.

Nobody watching wonders what they missed.

When does this not work well?

The steps have to be separable in space. An interface where the action happens across the whole screen at once, or a gesture that travels from top to bottom, cannot be framed tightly without losing one end of it.

How do you fix a result that came back wrong?

Say which section is wrong and which part of the screen it should have shown instead. Vizard Agent keeps the crop it chose for every section, so a correction changes one of them rather than re-cropping the entire recording from the beginning.

How does Vizard Agent compare to doing it yourself?

Automatic reframing follows the most active thing in the picture, and on a screen recording that is always the pointer rather than the interface. The result tracks perfectly and shows the viewer a cursor moving around instead of a tutorial, which is why Vizard Agent works from the steps instead.

Automatic reframing Vizard Agent
What it follows The pointer The step being performed
The crop Moves constantly Held per section
Gesture starts Frequently cropped out Checked for specifically
Between crops Jumps Eased
Changing format Re-tracks the cursor Sections re-derived

Common questions

Can it follow the cursor if I want? Yes, and Vizard Agent will show you why it usually reads badly.

How many crops will it use? As many as there are steps. Vizard Agent holds each one.

Can I give the sections myself? Yes. Give Vizard Agent rough times and which part of the screen.

Will the taps be visible? Vizard Agent checks the start of each gesture specifically.

What about square or vertical? Any shape. Vizard Agent re-derives the crops per format.

Will text stay readable? Cropping helps. Vizard Agent checks it at the delivered size.

Can it add tap indicators? Yes, if the recording does not show them; Vizard Agent adds them.

Does it speed up the waiting? Yes, if you ask. Say what counts as waiting.

Can it zoom rather than crop? Same thing here, and Vizard Agent eases the moves.

What if a step spans the whole screen? Vizard Agent will say so and leave that section wide.

Can it narrate the steps? Yes. Vizard Agent times the narration to the crops.

Will the recording get shorter? Only if you ask Vizard Agent to trim the waiting.

Can I see the crop plan first? Yes. Ask for the sections and their framing.

Why is following the cursor so bad? Because the pointer moves between actions as well as during them, so the frame is busiest exactly when nothing is happening.