Speaker diarization — who said what

Upload a recording and get a transcript with speaker diarization: every turn labeled by speaker, not one unbroken wall of text.

Drag a file here or browse

MP3, WAV, M4A, FLAC · MP4, MOV, MKV, WebM · up to 2 GB

Set language, translation, or vocabulary hints below before transcribing

OR

YouTube, Dropbox, Google Drive, or any direct MP3/MP4 link

Names, acronyms, or domain terms that improve accuracy

to turn on speaker labeling for this upload.

MP3, WAV, M4A, FLAC · MP4, MOV, MKV · up to 2 GB · 10 free minutes every day

What speaker diarization is

A raw transcript is just text — it does not say who spoke each line. Speaker diarization adds that layer: the recording is split into turns, and each turn is tagged with a speaker label, so a two-person conversation reads as a back-and-forth instead of a single paragraph.

convert.express supports speaker diarization on every transcription model, using two different techniques depending on which one you choose.

Two ways diarization runs

  • OpenAI Whisper and Microsoft MAI-Transcribe 1.5: a GPT-based diarization pass runs after transcription and labels each segment. This works across both models because it operates on the finished transcript, not inside either model.
  • Google Gemini 3.5 Transcribe: diarization happens in-model, as part of the same transcription request, for up to 8 speakers. No second pass is needed, and this is the model chosen when diarization matters most.

What the output looks like

This is a real, unedited excerpt from a convert.express job — Gemini 3.5 Transcribe's in-model diarization on a two-person conversation. Filler words and false starts are left exactly as the model heard them.

Speaker A: I think that a really big problem and this is the voice because
the voice is a very personal thing that you expose to the audience.

Speaker B: Part of your body basically.

Speaker A: Yes, you know part of your body or of your mind. So it's very
personal so but if you choose to do it in your life, if you decide to
to start in to be a singer

Speaker B: Right.

Speaker A: you you you should you should try to expose yourself also in
the social media because it's the same. So I think it could could be a
a beautiful tirocinio and this is pre-work. I don't know what to say.

Speaker B: Like beautiful try, beautiful experimental

Speaker A: Training, yeah training, beautiful training to work.

See the full job end to end — transcript, SRT, VTT, PDF, summary, action items and a translated SRT, all from this same recording.

What diarization does not do

Speakers come back as generic labels — Speaker 1, Speaker 2 — based only on what the model detects in that one recording. There is no UI to rename a speaker once and have it apply globally across future jobs; each transcript's labels are its own. If you need persistent named speakers across a series, that is not something convert.express does today.

Frequently asked questions

What is speaker diarization?

Speaker diarization is the process of splitting a transcript by who is talking, not just what was said. Instead of one unbroken block of text, you get each turn labeled — "Speaker 1", "Speaker 2" — so a conversation reads the way it happened.

How does convert.express do it?

Two ways, depending on the model. With OpenAI Whisper or Microsoft MAI-Transcribe 1.5, a separate GPT-based diarization pass runs after transcription and labels each segment. With Google Gemini 3.5 Transcribe, diarization happens in-model as part of the same request, for up to 8 speakers, with no second pass needed.

What does the output actually look like?

Each line of the transcript is prefixed with its speaker label, in order, with no manual sorting needed. See a real example below, pulled from an actual convert.express job.

Can I rename "Speaker 1" to an actual name?

Not yet, and not globally. Speakers come back labeled generically per job — "Speaker 1", "Speaker 2" — based on what the model detects in that recording. There is no relabeling UI, and a name typed for one job does not carry over to the next.

Is diarization free?

Yes — turning it on costs nothing extra beyond the usual per-minute transcription price, and the first 10 minutes of any day are free regardless of model.

Try speaker diarization

10 minutes free every day. No account required to start.

Transcribe a recording →

See also: subtitle export · model comparison · meeting summaries