convert.express

Transcribe Cantonese Audio or Video to Text

Upload any Cantonese recording — MAI-Transcribe 2 detects the language automatically and returns an accurate transcript in seconds.

Drag a file here or browse

MP3, WAV, M4A, FLAC · MP4, MOV, MKV, WebM · up to 2 GB

Set language, translation, or vocabulary hints below before transcribing

OR

YouTube, Dropbox, Google Drive, or any direct MP3/MP4 link

Removes fillers and false starts for readable text

Names, acronyms, or domain terms that improve accuracy

MP3, WAV, M4A, FLAC · MP4, MOV, MKV · up to 2 GB · 10 minutes free every day

粵語 · convert.express →

About Cantonese transcription

Spoken across Hong Kong, Macau, and Guangdong province, Cantonese sits within the Sino-Tibetan language family but is a fully distinct language from Mandarin — not a dialect of it — with its own phonology, vocabulary, and writing conventions. Where Mandarin uses four tones, Cantonese has six to nine phonemic tones depending on how you count them, and that density creates real ambiguity for any speech recognition system trying to map a spoken syllable to the correct character. Add the fact that Cantonese writing uses Chinese characters including Cantonese-specific ones absent from Mandarin texts, and the output layer becomes more complex than many other Chinese-language transcription tasks. Colloquial Cantonese speech also frequently mixes these native characters with Mandarin-standard forms in ways that do not follow a single stable convention. Transcription accuracy for Cantonese on convert.express is rated medium, reflecting this combination of tonal complexity and script variation rather than any single bottleneck.

MAI-Transcribe 2 and Gemini 3.5 Transcribe both support Cantonese, detected automatically from the audio — no manual language selection is needed. OpenAI Whisper does not support this language. Upload any Cantonese recording up to 2 GB — MP3, WAV, M4A, FLAC, MP4, MOV, and most common audio and video formats are accepted. Every account includes 10 free minutes of transcription per day; longer files are billed at €0.02 per minute with a €0.50 minimum per transaction. Transcripts download as plain text and are automatically deleted within 24 hours of upload.

Good accuracy

Available transcription models

OpenAI Whisper — not supportedMAI-Transcribe 2PreviewGemini 3.5 TranscribePreview

Compare models →

10 min

Free every day

€0.02

Per minute beyond quota

< 1 min

Typical turnaround

Tips for the best results

  • Use a quiet environment — background noise is the biggest source of transcription errors.
  • Speak at a natural pace; extremely fast speech reduces accuracy in any language.
  • MP3 or M4A files work great; uncompressed WAV is ideal for professional recordings.
  • Files up to 2 GB are supported — long recordings are automatically split and merged.

Frequently asked questions

How accurate is AI transcription for Cantonese?
Cantonese transcription on convert.express is rated medium accuracy. The core difficulty comes from the language’s tonal system — six to nine phonemic tones compared to Mandarin’s four — which increases the chance of a spoken syllable being mapped to the wrong character. Cantonese also uses characters specific to its own writing tradition, distinct from standard Mandarin orthography, so the character-selection problem is broader than for many other Chinese-language inputs. Recordings with overlapping speakers or heavy code-switching between Cantonese and English tend to see the sharpest accuracy drop.
How long does Cantonese transcription take?
Most Cantonese recordings are ready in under a minute. Processing time scales with file length — a 10-minute recording typically returns results in 20–40 seconds, and a one-hour recording in around 4–6 minutes. Longer files are automatically split, transcribed in parallel, and merged, so you never need to cut your audio before uploading.
Can I translate Cantonese audio to English or other languages?
Yes. Single-pass translation to English is a Whisper feature, and Whisper does not support Cantonese — MAI-Transcribe 2 and Gemini 3.5 Transcribe transcribe Cantonese audio to Cantonese text instead. Once the transcript is ready you can translate it into English — or any of convert.express's other supported languages — using the Translate action in your dashboard. It's powered by Claude AI and also works on the summary and action list, not just the raw transcript.
Which transcription model should I use for Cantonese?
MAI-Transcribe 2 and Gemini 3.5 Transcribe both support Cantonese. MAI-Transcribe 2 is in preview, with whole-file speaker labeling, word-level timestamps, and entity biasing for proper names and terminology. Gemini 3.5 Transcribe is in preview, has the widest language coverage, and labels speakers in-model. OpenAI Whisper does not support Cantonese. Try them on a short clip to see which output suits your recording.
What audio and video formats are supported?
MP3, WAV, M4A, FLAC, OGG, Opus, MP4, MOV, MKV, and most other common audio and video formats are accepted. Files up to 2 GB can be uploaded. The language is detected automatically from the audio — no manual language selection is needed.

← See all 65 supported languages

粵語 · convert.express

粵語係香港、澳門同廣東省嘅主要語言,屬漢藏語系,但同普通話係兩種截然不同嘅語言,唔係普通話嘅方言,有自己嘅音韻系統、詞彙同書寫習慣。普通話得四個聲調,粵語就有六至九個聲調(視乎點計),聲調密度咁高,令語音識別系統喺判斷某個音節對應邊個字時更容易出錯。加上粵語書寫會用到一啲普通話文本冇嘅粵語專用字,輸出層面比起好多其他中文轉錄任務都複雜得多。日常粵語亦好常將呢啲粵字同普通話標準字形混用,而且冇一個統一固定嘅規範。convert.express 對粵語嘅轉錄準確度評級為中等,反映嘅係聲調複雜性同字形多樣性組合帶嚟嘅挑戰,而唔係單一某個問題。MAI-Transcribe 2 同 Gemini 3.5 Transcribe 都支援粵語,系統會自動從音頻偵測語言,唔需要手動揀。OpenAI Whisper 唔支援粵語。粵語錄音最大可上傳 2 GB,接受 MP3、WAV、M4A、FLAC、MP4、MOV 等大多數常見音頻同視頻格式。每個帳戶每日有 10 分鐘免費轉錄;超出部分收費 €0.02 每分鐘,每次交易最低消費 €0.50。轉錄文本可下載為純文字,並會於上傳後 24 小時內自動刪除。

AI 轉錄粵語嘅準確度點樣?
convert.express 對粵語轉錄嘅準確度評級為中等。主要難度來自粵語嘅聲調系統——六至九個聲調,比普通話嘅四個多得多——令音節映射到錯誤漢字嘅機率更高。粵語仲有自己書寫傳統特有嘅字,同標準普通話正字法唔同,所以選字嘅問題比其他好多中文輸入都複雜。如果錄音有多人同時講嘢,或者粵英夾雜嘅情況較多,準確度下跌會最明顯。
粵語轉錄要幾耐?
大多數粵語錄音可以喺一分鐘內完成處理。處理時間視乎檔案長度而定——10 分鐘嘅錄音通常 20 至 40 秒就有結果,一小時嘅錄音大約需要 4 至 6 分鐘。較長嘅檔案會自動分段、並行轉錄再合併,你唔需要喺上傳前自己剪片。
可以將粵語音頻翻譯成英文或其他語言嗎?
可以。一步直接翻譯成英文係 Whisper 嘅功能,但 Whisper 唔支援粵語——MAI-Transcribe 2 同 Gemini 3.5 Transcribe 會將粵語音頻轉錄成粵語文字。轉錄完成後,你可以喺 dashboard 用翻譯功能將文字翻譯成英文,或者 convert.express 支援嘅其他語言。呢個功能由 Claude AI 驅動,唔單止適用於原始轉錄文本,摘要同行動清單都同樣支援。
轉錄粵語應該揀邊個模型?
MAI-Transcribe 2 同 Gemini 3.5 Transcribe 都支援粵語。MAI-Transcribe 2 目前係預覽版,內建說話人標記、字詞級時間戳,以及針對專有名詞同術語嘅實體偏置功能。Gemini 3.5 Transcribe 亦係預覽版,語言覆蓋範圍最廣,並由模型直接標記說話人。OpenAI Whisper 唔支援粵語。建議用一段短片段分別試試兩個模型,睇吓邊個輸出效果最適合你嘅錄音。
支援咩音頻同視頻格式?
接受 MP3、WAV、M4A、FLAC、OGG、Opus、MP4、MOV、MKV 等大多數常見音頻同視頻格式,最大可上傳 2 GB 嘅檔案。系統會自動從音頻偵測語言,唔需要手動揀。