Transcribe audio and video to text
Drop your file and get a transcript in under a minute. 10 free minutes every day, 64 languages, no subscription.
Drag a file here or browse
MP3, WAV, M4A, FLAC · MP4, MOV, MKV, WebM · up to 2 GB
Set language, translation, or vocabulary hints below before transcribing
YouTube, Dropbox, Google Drive, or any direct MP3/MP4 link
Not applied on Gemini 3.5 Transcribe — the model can't combine vocabulary hints with word-level timestamps
MP3, WAV, M4A, FLAC · MP4, MOV, MKV · up to 2 GB. 10 minutes free every day — longer files show a preview until you pay.
How it works
Upload your file
Drag and drop any audio or video file — MP3, WAV, M4A, FLAC, MP4, MOV. No account needed to get started.
AI transcribes it
Choose between OpenAI Whisper, Microsoft MAI-Transcribe 1.5, or Google Gemini 3.5 Transcribe. Language is detected automatically — 64 languages supported. Most recordings are ready in under a minute.
Free quota, then pay as you go
You get 10 free minutes of transcription every day. Files beyond that show a preview — pay only when you want the full result.
Generate more
From the finished job, request an AI summary, an action list, or a translation into any of 103 languages — applied straight to the completed transcript.
Simple, transparent pricing
No subscription, no monthly commitment — you pay only for what you transcribe.
- Choose your AI model — Whisper (57 languages), MAI-Transcribe 1.5 (fast, entity biasing), or Gemini 3.5 Transcribe (widest language coverage, in-model diarization)
- Language auto-detected — 64 languages supported
- Supports MP3, WAV, M4A, FLAC, MP4, MOV and more
- 10 free minutes of transcription every day
- Preview shown before you commit to paying
- Download as plain text
- Files automatically deleted after 24 hours
Example: a 30-minute interview = €0.60
Example: a 90-minute lecture = €1.80
Minimum charge per transaction: €0.50
About convert.express
convert.express is a browser-based transcription service that converts audio and video recordings to text using AI. Upload any common format — MP3, WAV, FLAC, M4A, OGG, MP4, MOV, or MKV — and receive an accurate transcript in under a minute. Language is detected automatically across 64 languages including English, Spanish, French, German, Chinese, Japanese, Arabic, and more.
Three AI models are available: OpenAI Whisper, covering 57 languages with strong multilingual accuracy; MAI-Transcribe 1.5, a preview model covering 43 languages that returns results faster and supports entity biasing for proper names and technical terms; and Gemini 3.5 Transcribe, a preview model covering 60 languages with in-model speaker diarization and up to 1,000 biasing terms. All three run on the same infrastructure and produce output in the language of the recording.
The first 10 minutes of audio are free every day — no credit card required to get started. Recordings beyond the daily quota are billed at €0.02 per minute with a minimum of €0.50 per transaction. Files are automatically deleted within 24 hours of upload and are never used to train AI models.
Got a specific file type? See format-specific guides for MP3, M4A, WAV, MP4, MOV, voice memos, and more →
Already have a transcript? Free browser tools for SRT, VTT, and plain-text conversion →