# convert.express > Pay-as-you-go audio & video transcription, powered by OpenAI Whisper, Microsoft MAI-Transcribe 2, and Google Gemini 3.5 Transcribe. Upload MP3/WAV/M4A/FLAC/MP4/MOV or paste a file URL — get a transcript as TXT, PDF, SRT or VTT, with optional speaker diarization, AI summaries, action items and translation. 10 free minutes daily, €0.02/minute after, €0.50 minimum charge, no subscription or account required to try. Files are deleted from convert.express storage within 24 hours and never used for AI training; MAI speaker labeling uses pyannoteAI temporary Media API storage for up to 48 hours. Separate browser-only subtitle tools at convert.express/tools never upload the file at all. ## Key facts - **Pricing**: 10 free minutes of transcription every day (no sign-up required). Beyond the free quota: €0.02 per minute of audio or video. Minimum charge per transaction: €0.50. Example: a 30-minute interview costs €0.60; a 90-minute lecture costs €1.80. - **How processing works**: Files are uploaded and transcribed on our servers, then deleted within 24 hours. The subtitle and text converters at convert.express/tools are different — those run entirely in your browser and the file is never uploaded. - **Privacy**: Uploaded files are automatically deleted 24 hours after your last activity on the job and are never used for AI model training. - **Languages**: 65 languages supported for transcription across three models (57 via OpenAI Whisper, 63 via MAI-Transcribe 2, 61 via Gemini 3.5 Transcribe, with 8 languages exclusive to MAI/Gemini: Assamese, Bengali, Cantonese, Gujarati, Malayalam, Odia, Punjabi, Telugu). Translation, a separate feature, covers 103 languages. - **Models**: OpenAI Whisper (~2.7% word error rate on clean benchmark English, rising to roughly 8% on real-world recordings with background noise, accents, or multiple speakers), Microsoft MAI-Transcribe 2 (vendor-reported improvement over its predecessor, ranking on the Artificial Analysis Word Error Rate leaderboard and the FLEURS benchmark, not independently measured by us), and Google Gemini 3.5 Transcribe (2.6% average WER, vendor-published, with in-model speaker diarization for up to 8 speakers) — different benchmarks, so not directly comparable to each other. These are published third-party and vendor benchmarks, not our own measurements, and vary by test set. - **Output formats**: TXT, PDF, SRT, and VTT from the same job. - **Speaker diarization**: Available on all three models — a separate GPT pass on Whisper, in-model on Gemini 3.5 Transcribe (up to 8 speakers), and a separate whole-file pyannote Precision-2 pass on MAI-Transcribe 2. Speaker labels are per-job and cannot currently be renamed globally. - **File limits**: Audio and video files up to 2 GB. Supported formats: MP3, WAV, M4A, AAC, OGG, FLAC, OPUS, WMA, AMR audio and MP4, MOV, MKV, WebM, AVI, M4V video. - **Turnaround**: Most recordings are ready in under a minute. - **Translation**: Two options. OpenAI Whisper can translate audio directly to English in a single pass at upload time, at no extra cost. Separately, once any job finishes on any model, the dashboard's Translate action uses Claude AI to translate the transcript, summary, action list, or SRT/VTT subtitles into any of 103 languages, not just English — translated subtitles keep the original timecodes, since translation reuses the stored subtitle segments rather than re-transcribing. - **Operating since**: 2026-04-01. ## Docs - [Pricing & how it works](https://convert.express/pay-as-you-go): Full breakdown of the free quota, per-minute rate, minimum charge, and billing model. - [Supported languages](https://convert.express/languages): Complete list of 65 transcription languages with per-model availability. - [Transcription models](https://convert.express/models): Comparison of OpenAI Whisper, MAI-Transcribe 2, and Gemini 3.5 Transcribe accuracy, speed, and language coverage. - [Use cases](https://convert.express/for): Who uses convert.express — journalists, researchers, students, podcasters, and more. - [Subtitle & caption tools](https://convert.express/tools): Free browser-based tools for converting and editing subtitle files (SRT, VTT) — these run client-side; transcription does not. - [FAQ](https://convert.express/faq): Common questions about pricing, privacy, file formats, languages, AI models, diarization, sharing, and getting started without an account. - [Private transcription](https://convert.express/privacy-first-transcription): How convert.express handles privacy — no account needed, 24-hour auto-delete, no AI training on user data. - [OpenAI Whisper transcription](https://convert.express/whisper-transcription): How Whisper works, accuracy, languages (57), single-pass translation to English at upload time, and when to choose it. - [MAI-Transcribe 2 transcription](https://convert.express/mai-transcription): How MAI-Transcribe 2 works, whole-file speaker diarization, 60 languages, entity biasing, transcript styles, and public preview caveats. - [Gemini 3.5 Transcribe](https://convert.express/gemini-transcription): How Gemini 3.5 Transcribe works, 61 languages, in-model speaker diarization for up to 8 speakers, and public preview caveats. - [About](https://convert.express/about): Who builds convert.express, why it's priced the way it is, what it runs on, and when it isn't the right tool. - [Privacy policy](https://convert.express/privacy): Data handling, retention, and deletion practices. - [Terms of service](https://convert.express/terms): Usage terms and conditions. ## Optional - [SRT to TXT converter](https://convert.express/tools/srt-to-txt): Strip subtitle timecodes and convert SRT files to plain text. - [SRT to VTT converter](https://convert.express/tools/srt-to-vtt): Convert SubRip (.srt) subtitle files to WebVTT (.vtt) format. - [VTT to SRT converter](https://convert.express/tools/vtt-to-srt): Convert WebVTT (.vtt) subtitle files to SubRip (.srt) format. - [SRT time-shift tool](https://convert.express/tools/srt-shift): Shift all timestamps in an SRT file forward or backward by a fixed offset.