convert.express

Convert WAV to text

Upload an uncompressed WAV recording and get a transcript back — no quality lost, no conversion needed.

Drag a file here or browse

MP3, WAV, M4A, FLAC · MP4, MOV, MKV, WebM · up to 2 GB

Set language, translation, or vocabulary hints below before transcribing

OR

YouTube, Dropbox, Google Drive, or any direct MP3/MP4 link

Not applied on Gemini 3.5 Transcribe — the model can't combine vocabulary hints with word-level timestamps

.wav accepted directly · up to 2 GB · 10 minutes free every day

How it works

1

Upload your file

Drag and drop your .wav file onto the box below, or paste a link. No account needed to start.

2

AI transcribes it

OpenAI Whisper, Microsoft MAI-Transcribe 1.5, or Google Gemini 3.5 Transcribe converts the audio to text, detecting the language automatically.

3

Download the transcript

Get plain text back, plus SRT and VTT subtitle files when the recording carries timing information.

4

Generate more

From the finished job, request an AI summary, an action list, or a translation into any of 103 languages — applied straight to the completed transcript, no re-upload needed.

About .wav files

WAV stores audio uncompressed: every sample is kept, nothing is thrown away to save space. That makes it the format of choice for professional field recorders, studio interfaces, and any workflow where the audio might be edited or remixed later and needs to survive multiple passes without degrading further.

The tradeoff is size. A stereo WAV at standard 44.1kHz sampling runs roughly 10MB per minute, so a two-hour interview recorded in stereo can approach 1.2GB before you've done anything with it — the fastest of any common format to reach the 2 GB upload ceiling. A three-hour stereo recording at CD quality will generally cross that line. If you're recording something that long, capturing it as mono (most interviews and lectures don't need stereo separation) roughly halves the file size, or export a compressed copy (M4A or MP3 at 128kbps+) for upload — compression doesn't measurably change speech transcription accuracy, only the uncompressed file's size.

Because WAV involves no lossy compression, there's no bitrate tradeoff to think about the way there is with MP3 or AAC — accuracy on a WAV file depends entirely on the original recording quality (microphone, room, distance from the speaker), not the file format.

Upload the file as recorded. You'll get a plain-text transcript back, plus SRT and VTT subtitle downloads if the recording carries timing information — no need to convert or compress the WAV first.

Where these files usually come from

  • Dedicated field recorders (Zoom H-series, Tascam)
  • Studio audio interfaces
  • Screen-capture software with a separate audio export
  • Lossless rips from other equipment

Things to know

Stereo WAV recorded at standard CD quality (44.1kHz) reaches roughly 10MB per minute, so long recordings — around 3 hours in stereo — hit the 2 GB upload ceiling before shorter recordings in any other format would.

Record in mono when stereo separation isn't needed (most single-speaker or interview recordings), or export a compressed copy (M4A or MP3) for the upload — speech transcription accuracy is unaffected by that conversion.

Simple pay-as-you-go pricing

10 minutes free every day. Beyond that, €0.02/minute with a €0.50 minimum — no subscription. Files are automatically deleted within 24 hours.

See full pricing →

Frequently asked questions

Does WAV transcribe more accurately than MP3?
Not meaningfully, for typical recordings. WAV avoids the lossy compression MP3 uses, but at normal MP3 bitrates (128kbps and up) the difference in transcription accuracy is negligible. WAV's main advantage is for editing workflows, not transcription.
Why is my WAV file so much bigger than an MP3 of the same recording?
WAV stores every audio sample uncompressed, which runs about 10MB per minute in stereo at standard quality. An MP3 of the same recording is typically 5–10× smaller because lossy compression strips out inaudible detail.
What happens if my WAV file is close to the 2 GB limit?
Re-export as mono if the recording is single-speaker or interview-style audio (this roughly halves the file size), or export a compressed M4A/MP3 copy for upload — that has no meaningful effect on transcription accuracy.
Can I get SRT or VTT files from a WAV upload?
Yes, alongside the plain-text transcript, whenever the recording carries timing data.
Is a WAV file kept private after transcription?
Yes. It is automatically deleted within 24 hours of upload and never used to train AI models.

Already have subtitles, not audio?

If you're trying to convert an existing SRT or VTT file rather than transcribe new audio, the browser-based subtitle tools handle that instantly, with nothing uploaded anywhere.

See subtitle tools →
← See all formats