Convert WAV to text
Upload an uncompressed WAV recording and get a transcript back — no quality lost, no conversion needed.
Drag a file here or browse
MP3, WAV, M4A, FLAC · MP4, MOV, MKV, WebM · up to 2 GB
Set language, translation, or vocabulary hints below before transcribing
YouTube, Dropbox, Google Drive, or any direct MP3/MP4 link
Not applied on Gemini 3.5 Transcribe — the model can't combine vocabulary hints with word-level timestamps
.wav accepted directly · up to 2 GB · 10 minutes free every day
How it works
Upload your file
Drag and drop your .wav file onto the box below, or paste a link. No account needed to start.
AI transcribes it
OpenAI Whisper, Microsoft MAI-Transcribe 1.5, or Google Gemini 3.5 Transcribe converts the audio to text, detecting the language automatically.
Download the transcript
Get plain text back, plus SRT and VTT subtitle files when the recording carries timing information.
Generate more
From the finished job, request an AI summary, an action list, or a translation into any of 103 languages — applied straight to the completed transcript, no re-upload needed.
About .wav files
WAV stores audio uncompressed: every sample is kept, nothing is thrown away to save space. That makes it the format of choice for professional field recorders, studio interfaces, and any workflow where the audio might be edited or remixed later and needs to survive multiple passes without degrading further.
The tradeoff is size. A stereo WAV at standard 44.1kHz sampling runs roughly 10MB per minute, so a two-hour interview recorded in stereo can approach 1.2GB before you've done anything with it — the fastest of any common format to reach the 2 GB upload ceiling. A three-hour stereo recording at CD quality will generally cross that line. If you're recording something that long, capturing it as mono (most interviews and lectures don't need stereo separation) roughly halves the file size, or export a compressed copy (M4A or MP3 at 128kbps+) for upload — compression doesn't measurably change speech transcription accuracy, only the uncompressed file's size.
Because WAV involves no lossy compression, there's no bitrate tradeoff to think about the way there is with MP3 or AAC — accuracy on a WAV file depends entirely on the original recording quality (microphone, room, distance from the speaker), not the file format.
Upload the file as recorded. You'll get a plain-text transcript back, plus SRT and VTT subtitle downloads if the recording carries timing information — no need to convert or compress the WAV first.
Where these files usually come from
- Dedicated field recorders (Zoom H-series, Tascam)
- Studio audio interfaces
- Screen-capture software with a separate audio export
- Lossless rips from other equipment
Things to know
Stereo WAV recorded at standard CD quality (44.1kHz) reaches roughly 10MB per minute, so long recordings — around 3 hours in stereo — hit the 2 GB upload ceiling before shorter recordings in any other format would.
Record in mono when stereo separation isn't needed (most single-speaker or interview recordings), or export a compressed copy (M4A or MP3) for the upload — speech transcription accuracy is unaffected by that conversion.
Simple pay-as-you-go pricing
10 minutes free every day. Beyond that, €0.02/minute with a €0.50 minimum — no subscription. Files are automatically deleted within 24 hours.
Frequently asked questions
- Does WAV transcribe more accurately than MP3?
- Not meaningfully, for typical recordings. WAV avoids the lossy compression MP3 uses, but at normal MP3 bitrates (128kbps and up) the difference in transcription accuracy is negligible. WAV's main advantage is for editing workflows, not transcription.
- Why is my WAV file so much bigger than an MP3 of the same recording?
- WAV stores every audio sample uncompressed, which runs about 10MB per minute in stereo at standard quality. An MP3 of the same recording is typically 5–10× smaller because lossy compression strips out inaudible detail.
- What happens if my WAV file is close to the 2 GB limit?
- Re-export as mono if the recording is single-speaker or interview-style audio (this roughly halves the file size), or export a compressed M4A/MP3 copy for upload — that has no meaningful effect on transcription accuracy.
- Can I get SRT or VTT files from a WAV upload?
- Yes, alongside the plain-text transcript, whenever the recording carries timing data.
- Is a WAV file kept private after transcription?
- Yes. It is automatically deleted within 24 hours of upload and never used to train AI models.
Related formats
Built for your workflow
Transcription for Researchers
Fast, accurate transcription for interviews, focus groups, and fieldwork — supporting 64 languages and multiple files at once.
Transcription for Lawyers
Fast, private transcription for depositions, client calls, and hearings — audio or video, with automatic 24-hour file deletion.
64 languages supported — language is detected automatically.
Already have subtitles, not audio?
If you're trying to convert an existing SRT or VTT file rather than transcribe new audio, the browser-based subtitle tools handle that instantly, with nothing uploaded anywhere.
See subtitle tools →