Vizard Agent

How to talk over a screen recording to direct the edit

Last updated 2026-08-30 · 7 min read

Record your screen while talking through what matters, then tell Vizard Agent the audio is a brief rather than a soundtrack. Vizard Agent transcribes what you said, uses it to find and name each section, builds a contact sheet per part, and delivers the silent cut you described out loud.

What is the short version?

Writing an edit brief is slow and describing an interface in words is worse. Talking over the recording while you use the app is faster than either, and the words can be thrown away once they have done their job.

  1. Go to Vizard Agent and upload the screen recording.
  2. Say the spoken track is a guide for editing, not for the output.
  3. Say exactly what you want delivered.

What do you need before you start?

The recording and a spoken commentary. Vizard Agent transcribes what you said to work out which parts of the recording matter, so you can be as informal as you like — the words never appear in the finished video and nobody else hears them.

What do you type into Vizard Agent?

State plainly that the audio is instructions. Vizard Agent will otherwise treat your voice as the video's soundtrack, which is the reasonable reading and is precisely the opposite of what you meant when you were talking yourself through the interface.

Prompt

ONE CLIP ONLY — no extra clips, variants or recuts. This is a silent screen recording of [the app]. I speak [Polish] on the audio track; that audio is a GUIDE for editing only and must not appear in the output.

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real recording where the narration was a brief. The structure comes from the speech: it transcribed the commentary, then organised the footage into the sections that commentary named.

  1. Inspects the source video's dimensions, duration and stream properties.
  2. Transcribes the shots and the speech to identify the key sections.
  3. Reads the beginning of the spoken transcript to understand the intent.
  4. Analyses the shot and transcript data together.
  5. Extracts key frames and generates contact sheets for inspection.
  6. Reviews the first batch of frames — the home screen, the catalogue, the search.
  7. Reviews the second batch — the card details and the collection.
  8. Reviews the third batch — the profile, the community map and the rest.
  9. Checks the final seconds, extracts granular frame sequences for each section, and builds a separate contact sheet per story beat.

Step 9 is where the spoken brief pays off. Because the commentary named the sections, the footage can be grouped by what the app was doing rather than by the clock, and each part of the interface gets reviewed as a unit.

What does the result look like?

The recording this was built from measures 1080x1920, H.264, 30fps, 275.13 seconds — four and a half minutes of app usage with a running commentary. From it comes exactly one clip, silent, organised into the sections the commentary described, and nothing else.

The instruction to deliver one thing matters more than it sounds. Left open, a request like this reasonably produces a main cut plus some shorts, and if you only wanted one file that is four extra things to look through.

When does this not work well?

Speaking a brief is much faster than writing one and correspondingly less precise, and some of that imprecision inevitably reaches the edit. These are the places where Vizard Agent can only be as clear as the commentary was.

How do you fix a result that came back wrong?

Point at the section using the name you gave it out loud on the recording. Vizard Agent keeps the transcript, all the per-section contact sheets and the frame sequences, so one part can be re-cut without re-reading the recording.

How does Vizard Agent compare to doing it yourself?

By hand this means writing a shot list for an interface, which is tedious and error-prone, or editing it yourself while remembering what you meant to show. Talking through it once is faster than both and is what most people do anyway when explaining it to somebody else.

By hand Vizard Agent
The brief Written, slowly Spoken while recording
Structure By timecode By the sections you named
Review Scrub the whole file A contact sheet per section
Deliverables Whatever seems useful Exactly what you asked for

Common questions

Does my voice end up in the video? Not if you say it is a guide. Vizard Agent uses it to edit and then discards it.

Can I speak my own language? Yes. Vizard Agent transcribes it regardless of the output language.

How detailed should the commentary be? Name the sections and say what matters in each. That is enough to structure the edit.

Can I ask for exactly one file? Yes, and it is worth saying. Vizard Agent otherwise delivers useful extras you did not want.

Does it work for anything other than apps? Yes. Vizard Agent handles any recording you can talk over — a workshop, a spreadsheet, a game.

Why speak the brief rather than write it? Because describing an interface in text is slow and imprecise, whereas talking while you use it produces a description that is already synchronised to the footage it refers to.

What if I ramble? Vizard Agent works from whatever you named. Rambling produces a looser structure, not a broken one.

Can it add narration afterwards? Yes, written properly and recorded separately, once the structure is settled.

Does the recording need to be tidy? No. Your commentary is what tells Vizard Agent which parts of it were the point.

Can I record a second pass of commentary? Yes, and it often helps. Vizard Agent takes a fresh spoken pass over the same footage as a revised brief, which is quicker than writing corrections out afterwards.

Is this quicker than writing a brief? Considerably, for anything visual. Vizard Agent gets a description already synchronised to the footage, which is something no written brief about an interface can give it.

What should I say while recording? Name each section as you reach it and say why it matters. Vizard Agent uses those names as the structure, so a running label is worth more than a running commentary.

Does Vizard Agent check the result? Yes. It reviews the sections against the transcript before delivery.