Vizard Agent

How to turn a raw consultation recording into a case study video

Last updated 2026-08-18 · 6 min read

Upload the raw session to Vizard Agent and say what the case study should show. Vizard Agent transcribes it, measures the room's noise floor and looks at the spectrum before choosing a cleanup, tests crop framings at full size and compares them, then cuts the recording into a narrative rather than a trimmed transcript.

What is the short version?

A recorded consultation is the most valuable footage a practice has and the least usable. The room is noisy, the framing is wrong for a feed, and the story is buried in forty minutes of talking that nobody outside the room will sit through.

  1. Go to Vizard Agent and upload the raw recording.
  2. Say what the case study demonstrates and who it is for.
  3. Say the length, the shape and how much cleanup the audio needs.

What do you need before you start?

The recording and the story you want it to tell. Vizard Agent measures the audio and tests the framing itself, so what only you can supply is the argument the footage is being cut to make, since a consultation contains several possible stories and only one of them is the case study.

What do you type into Vizard Agent?

Describe the finished piece, not the edit. Vizard Agent reads a brief written for a human editor as instructions, so a paragraph naming the story, the cleanup and the format gets you further than a shot list, and Vizard Agent will handle the sequencing from the transcript.

Prompt

Edit this raw consultation into a [4] minute narrative case study for [Instagram], vertical. Clean up the audio. [What the case shows.]

Variants worth knowing:

What does Vizard Agent actually do?

Here is the order Vizard Agent worked in on a real case study. Two stretches stand out: it looked at a picture of the noise before deciding how to clean it, and it rendered both crop options rather than choosing one from a thumbnail.

  1. Probes the recording and reviews the project's history.
  2. Extracts sample frames and transcribes the consultation audio.
  3. Measures the framing with a coordinate grid and inspects the overlay to plan the crop.
  4. Reads the transcript in full, in three passes, and checks the word-level timings.
  5. Downloads the footage and measures the audio levels, the overall loudness and the silent regions.
  6. Analyses the noise floor and the spectrum, then looks at the spectrum image.
  7. Verifies the key lines and checks the practitioner's name rather than repeating it blind.
  8. Tests crop windows across the whole video and reviews the candidates.
  9. Renders both framings at full size, compares the standard and the close-up side by side, then tests the noise cleanup chain and measures the speaker's level before and after.

Step 6 is the one nobody expects. Choosing a noise reduction by ear on a laptop is guesswork; looking at where the noise actually sits in the spectrum tells Vizard Agent what to remove and, more importantly, what to leave alone.

What does the result look like?

From the run this page is written from, probed on the delivered file: 1080x1920, H.264, 30fps, 243.9 seconds, AAC audio. Vertical, just over four minutes, reframed from a wide recording, with the room noise reduced and the speech levelled.

Four minutes is unusually long for a vertical video and correct for this format. A case study is watched by people evaluating a service, and they will sit through length that a casual scroller would not, so Vizard Agent paces it for attention rather than for the first three seconds.

When does this not work well?

A consultation is a private conversation before it is footage. Vizard Agent cleans it, cuts it and frames it, and every question about whether it should be published at all sits outside what the tool can judge, which in this format is the part that actually decides whether the video ships.

How do you fix a result that came back wrong?

Say which part. Vizard Agent keeps the transcript, the word timings, the noise analysis and both rendered framings, so changing the crop or easing the cleanup is a re-render against measurements it already made rather than a fresh pass over four minutes.

How does Vizard Agent compare to doing it yourself?

By hand this is watching forty minutes twice, marking the usable lines, guessing at a noise reduction, cropping to vertical and discovering at minute three that the subject drifts out of frame. Vizard Agent front-loads the measuring so those discoveries happen before the cut, not after it.

By hand Vizard Agent
Finding the story Watch it repeatedly Read from the full transcript
The noise Apply a preset by ear Spectrum analysed, then chosen
The crop Set it once Tested across the whole video, both rendered
Judging the cleanup Listen and hope Levels measured before and after

Common questions

How long should a case study be? Three to five minutes. The audience is evaluating you, so they will watch longer than a casual viewer.

Can it remove the client's name? Yes. Say so and Vizard Agent cuts around it and keeps it out of the captions.

Will the audio actually be usable? Vizard Agent measures the noise and the speech level rather than applying a preset, which is what keeps the voice intact.

Can I get a short version for a feed? Yes. Ask for both and the short cut comes from the same transcript and analysis.

What if two people are on camera? A vertical crop fits one. Say which, or take the wide version instead.

Does it need the whole recording? Yes, ideally. Vizard Agent finds the story by reading all of it rather than a section you pre-selected.

Can it add captions? Yes, and for this audience it is usually worth it. Vizard Agent cuts them from the same word-level transcript it used to build the edit, so they land on the syllable rather than drifting.

Can I reuse the analysis for a second case study? Within the same project, yes. Vizard Agent keeps the transcript and the noise profile, so a second cut from the same session starts from the measurements rather than from the raw file again.