Gemini 3.5 Transcribe transcription online
Google's multilingual ASR model. 60+ languages · in-model speaker labels for up to 8 speakers · up to 1,000 biasing terms.
Drag a file here or browse
MP3, WAV, M4A, FLAC · MP4, MOV, MKV, WebM · up to 2 GB
Set language, translation, or vocabulary hints below before transcribing
YouTube, Dropbox, Google Drive, or any direct MP3/MP4 link
Not applied on Gemini 3.5 Transcribe — the model can't combine vocabulary hints with word-level timestamps
MP3, WAV, M4A, FLAC · MP4, MOV, MKV · up to 2 GB · 10 free minutes every day
What is Gemini 3.5 Transcribe?
Gemini 3.5 Transcribe is Google's automatic speech recognition model, published at a 2.6% word error rate. It runs in verbatim mode, which unlocks its two standout features: word-level timestamps and in-model speaker diarization for up to 8 speakers, with no separate diarization pass required.
2.6% WER is Google's own published figure, not our measurement — we do not have a ground-truth reference transcript to compute an independent WER, so we do not put one next to it. It also reflects Google's own test conditions, not the real-world figure quoted for Whisper elsewhere on this site.
The model's standout feature is language breadth: 60+ languages with automatic detection and mid-sentence code-switching, the widest coverage of the three models convert.express offers. It also accepts up to 1,000 custom-vocabulary terms to bias transcription towards names, product terms, or domain-specific vocabulary.
In our own testing, Gemini 3.5 Transcribe processed audio at roughly 150 seconds per hour of recording — slower than Whisper's 60–90 seconds and well behind MAI-Transcribe 1.5's under 15 seconds. Choose it for coverage and speaker labels, not for turnaround time.
Gemini 3.5 Transcribe at a glance
2.6%
WER (vendor-published)
60+
Languages
8
Speakers, in-model
1,000
Biasing terms
Gemini 3.5 Transcribe vs. Whisper vs. MAI-Transcribe 1.5
All three models cost €0.02/min. Full comparison →
| Gemini 3.5 Transcribe | Whisper | MAI-Transcribe 1.5 | |
|---|---|---|---|
| Word error rate | 2.6% (vendor-published) | ~2.7% clean / ~8% real-world | 2.4% (AA-WER leaderboard) |
| Languages | 60+ | 57 | 43 |
| Speed (our measurement, 1 hr audio) | ~150 s | 60–90 s | <15 s |
| Single-pass translation to English (upload time) | Not supported | Yes | Not supported |
| Speaker labeling | In-model, up to 8 | Separate pass | Separate pass |
| Entity/vocabulary biasing | Up to 1,000 terms | Basic hints | Up to 200 phrases |
| Status | Public Preview | Generally available | Public Preview |
| Price | €0.02/min | €0.02/min | €0.02/min |
All three models' transcripts can also be translated afterward into any supported language using the dashboard's Translate action — the row above only covers Whisper's free, single-pass translation option at upload time.
Gemini 3.5 Transcribe's WER figure is Google's vendor-published number; MAI-Transcribe 1.5's comes from the Artificial Analysis (AA-WER) leaderboard; Whisper's comes from published clean-benchmark and real-world comparisons. Three different benchmarks, three different conditions — not a strict apples-to-apples comparison. The speed row is our own measurement for all three models, run on identical test audio.
Honest limitations
- — Public preview, no SLA. If the model is unavailable, the job falls through to the next available provider.
- — Slower than Whisper and MAI-Transcribe 1.5 — roughly 150 seconds per hour of audio in our own testing, against 60–90 seconds for Whisper and under 15 for MAI-Transcribe 1.5.
- — No single-pass translation at upload time — Gemini 3.5 Transcribe is transcribe-only. Its transcripts can still be translated afterward, into English or any other supported language, via the dashboard's Translate action.
- — Preview model updates — the model may be updated without notice, which could change output characteristics.
Need the fastest turnaround? MAI-Transcribe 1.5 processes audio in under 15 seconds per hour.
Who benefits most from Gemini 3.5 Transcribe
Gemini 3.5 Transcribe is the right choice when language coverage or built-in speaker labels matter more than turnaround time. Its 60+ languages and in-model diarization are the main differentiators for multilingual and multi-speaker workflows.
- Researchers — in-model speaker labels for multi-participant interviews without a separate diarization pass.
- Journalists — the widest language coverage of the three models, for interviews conducted outside English.
- Students — transcribe lectures in languages Whisper and MAI-Transcribe don’t cover.
- Podcasters — automatic speaker labels for multi-guest episodes recorded in one file.
Frequently asked questions
What is Gemini 3.5 Transcribe?
- Gemini 3.5 Transcribe is Google's speech-to-text model, published at a 2.6% word error rate. It transcribes in verbatim mode with word-level timestamps and in-model speaker diarization for up to 8 speakers, and supports 60 languages with automatic detection and code-switching. convert.express offers it as a premium model option alongside OpenAI Whisper and MAI-Transcribe 1.5.
How accurate is Gemini 3.5 Transcribe compared to Whisper and MAI-Transcribe?
- Gemini 3.5 Transcribe is published at 2.6% word error rate by Google. Whisper is published at roughly 2.7% WER on clean benchmark English, rising to roughly 8% on real-world recordings; MAI-Transcribe 1.5 is published at 2.4% on the Artificial Analysis (AA-WER) leaderboard. These are three different benchmarks under three different test conditions, so none of the figures are a strict apples-to-apples comparison. We measured Gemini 3.5 Transcribe ourselves at roughly 150 seconds per hour of audio on our own test files — slower than both Whisper and MAI-Transcribe — but did not attempt to compute our own WER, since doing so honestly requires a ground-truth reference transcript we do not have.
How does speaker labeling work with Gemini 3.5 Transcribe?
- Unlike Whisper and MAI-Transcribe 1.5, which both identify speakers with a separate diarization pass after transcription, Gemini 3.5 Transcribe labels speakers in-model as part of the same request — up to 8 speakers (3 or more is experimental). This only runs in verbatim mode, which also produces the word-level timestamps convert.express uses to build subtitle cues.
What is custom vocabulary and how do I use it?
- Custom vocabulary (also called entity biasing) lets you supply a list of words and phrases that the model should expect to encounter, up to 1,000 terms accepted by the API — though convert.express caps submissions at 100, since Google's own documentation notes results degrade past that point. Use it for interviewee names, organisation names, product names, or technical terminology. Enter the phrases in the transcription settings before uploading your file.
What does "Public Preview" mean for Gemini 3.5 Transcribe?
- Gemini 3.5 Transcribe is in public preview on convert.express, which means it is available but carries no uptime SLA. If the model is unavailable, the job falls through to the next available provider and the failure reason is recorded on the job. Preview status also means the model may change behaviour without notice.
When should I choose Gemini 3.5 Transcribe over Whisper or MAI-Transcribe?
- Choose Gemini 3.5 Transcribe when you need the widest language coverage (60 languages with auto-detection and code-switching), in-model speaker labels without a second pass, or the largest custom-vocabulary list of the three models. Choose Whisper when you need a free single-pass English translation at upload time or a model with a formal SLA. Choose MAI-Transcribe 1.5 when raw speed matters most. Either way, once a job finishes you can translate its transcript into English or any other supported language from your dashboard — that option isn't limited to Whisper.
Try Gemini 3.5 Transcribe free
10 minutes free every day. No account required.
Transcribe with Gemini 3.5 Transcribe →See also: Whisper transcription · MAI-Transcribe 1.5 · compare models · pricing