How to transcribe only the dialogue and skip the singing
Tell Vizard Agent that the sung passages are not dialogue. Speech recognition transcribes singing as readily as speech, so without that instruction a theme song and a lyric over a montage end up in the transcript looking exactly like lines of script.
What is the short version?
You want the spoken lines from a drama episode, and the episode has a theme song, an insert song over a montage and characters who occasionally sing. All of it lands in the transcript. Vizard Agent separates the two before transcribing rather than filtering afterwards.
- Tell Vizard Agent to transcribe dialogue only.
- Let it identify the sung sections first.
- Check the boundaries where speech turns into song.
What do you need before you start?
The video, and a decision about the edge cases. A character speaking over background music is dialogue; a character singing is not; a voice-over during an instrumental passage usually is. Those are your calls, and saying them up front saves a pass.
You should also say what you want done with the songs rather than leaving it implied. Dropping them silently and marking them as "[song]" produce very different documents, and if you are going to translate or subtitle from the transcript afterwards, the marker is usually what you want.
What do you type into Vizard Agent?
State the exclusion as part of the job, not as a correction afterwards. This is a case where the default behaviour is reasonable and simply not what you want, so the instruction has to arrive with the file.
Transcribe the spoken dialogue only. Leave out the singing, but mark where the songs are.
In the session behind this article the request came in later, after a first transcript had included everything, and the whole second pass existed only to separate the two. Saying it first would have cost nothing.
What does Vizard Agent actually do?
It classifies before it transcribes. The full audio gets pulled out and examined to work out which stretches are spoken and which are sung, and only then does the transcription run, against the spoken stretches. That order matters because deleting sung lines from a finished transcript means recognising them as sung from the text alone, which is unreliable.
In that session the working sequence was:
- Extracted the complete audio track in order to classify it rather than working from the video.
- Checked how spoken passages differ from sung ones in that particular episode.
- Identified the speaking and the singing sections across the whole runtime.
- Cross-checked the on-screen Chinese subtitles at the dialogue stretches only.
- Zoomed into frames to read small subtitle lines where the audio was hard to make out.
- Checked the closing section separately so the sign-off lines were not dropped with the end theme.
That last step is the one that catches errors. Endings are where speech and song overlap most, and a classifier that gets the boundary wrong there loses real dialogue.
What does the result look like?
A transcript of the spoken lines with timestamps, and markers where the songs were rather than silent gaps. It reads as a script, which is the point — a transcript with the theme song lyrics in the middle of it is not usable as a script or as a translation source.
Vizard Agent cross-checks against the on-screen subtitles where they exist, so lines that the audio renders ambiguously get resolved against text that is already correct. On archaic or formal dialogue that is often the difference between a usable transcript and a plausible-looking wrong one.
When does this not work well?
When someone speaks over the song. Dialogue under a music bed is common and the classification has to decide whether the stretch is speech-with-music or song, and it will sometimes get it wrong in both directions. Vizard Agent flags the ambiguous stretches rather than silently picking.
Half-sung delivery is genuinely hard. Chanting, recitative, a rap verse and a character speaking rhythmically all sit between the two categories, and there is no correct answer that does not depend on what you want the transcript for. Say which way you want those handled.
And it does not help if the audio is a single flattened mix with the song at speech level throughout. Then everything is speech-with-music and the classification has nothing to separate. A transcript with markers you place by hand is the more honest outcome.
How do you fix a result that came back wrong?
Name the timestamp and say which way it went wrong. "There is dialogue missing around 14:20" and "the chorus is still in at 3:40" are different errors with different fixes, and Vizard Agent re-examines that stretch rather than re-running the whole classification.
If whole sections came out wrong, the boundary rule is probably the issue rather than any individual call. Restate what counts as dialogue — especially whether speech over music is in or out — and Vizard Agent reclassifies against that rule.
If the words are right but the language is wrong, that is a separate step. Vizard Agent transcribes in the spoken language first and translates afterwards, so a translation problem does not mean the transcript was wrong.
How does Vizard Agent compare to doing it yourself?
Every transcription tool will hand you the songs along with the speech, and the usual fix is to read the transcript and delete what looks like lyrics. That works on a theme song with an obvious chorus and fails on an insert song that sounds like a monologue in text.
The difference is working from the audio rather than from the text. Vizard Agent decides what is sung by listening to it, where you would be deciding from the transcript after the distinguishing information has already been thrown away. On a drama episode with several songs in it, that is the difference between a clean script and an afternoon of judgement calls.
Common questions
Will the timestamps still be right? Yes. Vizard Agent keeps the original timings, so removed songs leave gaps rather than shifting everything.
Can I get the lyrics separately? Yes. Ask and Vizard Agent transcribes the sung sections into their own file.
What about dialogue over background music? That counts as dialogue by default. Tell Vizard Agent if you want it excluded.
Does this work on any language? Yes. The classification is on the audio rather than on the words.
Can it handle several episodes? Yes. Vizard Agent applies the same rule across a stack of files.
What if a character sings one line? It is marked as sung. Say if you would rather have short sung lines kept in.
Will it mark where the songs were? If you ask. Vizard Agent can leave markers or drop them entirely.
Can I get subtitles from this? Yes. A dialogue-only transcript is the right input for a subtitle file.
What about the on-screen subtitles? Vizard Agent reads them where they exist and uses them to resolve unclear audio.
Does it separate speakers too? It can, as a separate instruction. Say so and Vizard Agent labels the lines.
What if the song is in another language? It is still classified as song. Vizard Agent does not need to understand it to exclude it.
Can I see the section boundaries? Yes. Ask Vizard Agent for the speech and song ranges as a list.
Does the video need to be processed first? No. Vizard Agent extracts the audio itself.
How accurate is the split? Good in the middles, less certain at boundaries, which is why Vizard Agent checks those specifically.