convert.express

Transcribe Persian Audio or Video to Text

Upload any Persian recording — OpenAI Whisper detects the language automatically and returns an accurate transcript in seconds.

Drag a file here or browse

MP3, WAV, M4A, FLAC · MP4, MOV, MKV, WebM · up to 2 GB

Set language, translation, or vocabulary hints below before transcribing

OR

YouTube, Dropbox, Google Drive, or any direct MP3/MP4 link

Not applied on Gemini 3.5 Transcribe — the model can't combine vocabulary hints with word-level timestamps

MP3, WAV, M4A, FLAC · MP4, MOV, MKV · up to 2 GB · 10 minutes free every day

فارسی · convert.express →

About Persian transcription

Persian sits within the Iranian branch of the Indo-European family, which gives it a relatively regular phonological structure compared to many of its geographic neighbors. That regularity helps: spoken Farsi has consistent consonant-vowel patterns and no lexical tone, so a speech recognition system doesn't have to resolve pitch ambiguity the way it would for a tonal language. The Persian alphabet, a variant of Arabic script, omits short vowels in standard writing, but that's a challenge for text rendering rather than for audio-to-text conversion, where the acoustic signal carries the vowel information directly. The upshot is that standard spoken Iranian Farsi transcribes with high accuracy. The one area worth knowing about is the gap between Iranian Farsi and Afghan Dari. They share virtually all their grammar, but vocabulary and pronunciation diverge enough that a system trained narrowly on Tehran speech would stumble on Kabul speech. Convert.express covers both varieties, so researchers, journalists, and legal teams working with Dari recordings don't need a separate workflow.

OpenAI Whisper and Gemini 3.5 Transcribe both support Persian, detected automatically from the audio — no manual language selection is needed. MAI-Transcribe 1.5 does not support this language. Upload any Persian recording up to 2 GB — MP3, WAV, M4A, FLAC, MP4, MOV, and most common audio and video formats are accepted. Every account includes 10 free minutes of transcription per day; longer files are billed at €0.02 per minute with a €0.50 minimum per transaction. Transcripts download as plain text and are automatically deleted within 24 hours of upload.

High accuracy

Available transcription models

OpenAI WhisperMAI-Transcribe 1.5 — not supportedGemini 3.5 TranscribePreview

Compare models →

10 min

Free every day

€0.02

Per minute beyond quota

< 1 min

Typical turnaround

Tips for the best results

  • Use a quiet environment — background noise is the biggest source of transcription errors.
  • Speak at a natural pace; extremely fast speech reduces accuracy in any language.
  • MP3 or M4A files work great; uncompressed WAV is ideal for professional recordings.
  • Files up to 2 GB are supported — long recordings are automatically split and merged.

Frequently asked questions

How accurate is AI transcription for Persian?
Accuracy for Persian is high on convert.express, and the linguistic profile explains why. Iranian Farsi has no lexical tone and a regular sound system, which removes two common sources of transcription error. The Persian alphabet's omission of short vowels is a writing convention that doesn't affect audio processing. The bigger variable is dialect: Iranian Farsi and Afghan Dari diverge in vocabulary and pronunciation despite sharing the same grammar, and the service covers both, so you can expect consistent results whether your audio comes from Tehran or Kabul.
How long does Persian transcription take?
Most Persian recordings are ready in under a minute. Processing time scales with file length — a 10-minute recording typically returns results in 20–40 seconds, and a one-hour recording in around 4–6 minutes. Longer files are automatically split, transcribed in parallel, and merged, so you never need to cut your audio before uploading.
Can I translate Persian audio to English or other languages?
Yes, two ways. For a quick English-only result, enable "Translate to English" in the upload options above — OpenAI Whisper transcribes and translates Persian audio to English text in a single pass, at no extra cost. For any other target language, first transcribe normally, then use the Translate action on the finished transcript in your dashboard — it's powered by Claude AI and can translate the Persian transcript (or its summary and action list) into any of convert.express's other supported languages, not just English.
Which transcription model should I use for Persian?
OpenAI Whisper and Gemini 3.5 Transcribe both support Persian. OpenAI Whisper is the established choice, with broad language coverage and single-pass translation to English. Gemini 3.5 Transcribe is in preview, has the widest language coverage, and labels speakers in-model. MAI-Transcribe 1.5 does not support Persian. Try them on a short clip to see which output suits your recording.
What audio and video formats are supported?
MP3, WAV, M4A, FLAC, OGG, Opus, MP4, MOV, MKV, and most other common audio and video formats are accepted. Files up to 2 GB can be uploaded. The language is detected automatically from the audio — no manual language selection is needed.

← See all 64 supported languages

فارسی · convert.express

زبان فارسی شاخه‌ای از خانواده هند-اروپایی است و در مقایسه با بسیاری از زبان‌های همسایه، ساختار آوایی نسبتاً منظمی دارد. این ویژگی به نفع تشخیص گفتار است: فارسی محاوره‌ای الگوهای همخوان-واکه‌ای یکدستی دارد و فاقد تکیه واژگانی است، بنابراین سیستم تشخیص گفتار نیازی به رفع ابهام زیروبمی ندارد — برخلاف زبان‌های تکیه‌دار که این کار را پیچیده می‌کنند. الفبای فارسی، که گونه‌ای از خط عربی است، در نوشتار استاندارد مصوت‌های کوتاه را حذف می‌کند؛ اما این موضوع چالش رندرینگ متن است، نه تبدیل صدا به نوشتار — چون در پردازش صوتی، اطلاعات مصوت مستقیماً از سیگنال صوتی به دست می‌آید. نتیجه این است که فارسی گفتاری استاندارد ایرانی با دقت بالایی رونویسی می‌شود. یک نکته مهم که باید بدانید تفاوت میان فارسی ایرانی و دری افغانستانی است. این دو زبان تقریباً دستور زبان مشترکی دارند، اما واژگان و تلفظشان به اندازه‌ای با هم متفاوت است که سیستمی که صرفاً بر اساس گفتار تهرانی آموزش دیده، در برابر گفتار کابلی دچار مشکل می‌شود. convert.express هر دو گونه را پشتیبانی می‌کند، بنابراین محققان، روزنامه‌نگاران و تیم‌های حقوقی که با ضبط‌های دری کار می‌کنند نیازی به فرآیند جداگانه‌ای ندارند. OpenAI Whisper و Gemini 3.5 Transcribe هر دو از زبان فارسی پشتیبانی می‌کنند و زبان به‌صورت خودکار از روی صدا تشخیص داده می‌شود — نیازی به انتخاب دستی زبان نیست. MAI-Transcribe 1.5 این زبان را پشتیبانی نمی‌کند. هر فایل صوتی فارسی تا حجم 2 GB را می‌توانید آپلود کنید — فرمت‌های MP3، WAV، M4A، FLAC، MP4، MOV و اکثر فرمت‌های رایج صوتی و تصویری پذیرفته می‌شوند. هر حساب کاربری روزانه 10 دقیقه رونویسی رایگان دارد؛ برای فایل‌های طولانی‌تر، نرخ €0.02 به ازای هر دقیقه اعمال می‌شود و حداقل مبلغ هر تراکنش €0.50 است. فایل‌های متنی رونویسی‌شده قابل دانلود هستند و ظرف 24 ساعت پس از آپلود به‌صورت خودکار حذف می‌شوند.

دقت رونویسی هوش مصنوعی برای زبان فارسی چقدر است؟
دقت رونویسی فارسی در convert.express بالاست و ویژگی‌های زبانی این زبان دلیل آن را روشن می‌کنند. فارسی ایرانی نه تکیه واژگانی دارد و نه دستگاه آوایی پیچیده‌ای، که همین دو عامل از رایج‌ترین منابع خطا در رونویسی هستند. حذف مصوت‌های کوتاه در الفبای فارسی یک قرارداد نوشتاری است و تأثیری بر پردازش صوتی ندارد. متغیر مهم‌تر گویش است: فارسی ایرانی و دری افغانستانی با وجود دستور زبان مشترک، در واژگان و تلفظ تفاوت‌هایی دارند، و این سرویس هر دو را پوشش می‌دهد — بنابراین چه صدایتان از تهران باشد چه از کابل، می‌توانید انتظار نتایج یکسانی داشته باشید.
رونویسی فارسی چقدر طول می‌کشد؟
اکثر ضبط‌های فارسی در کمتر از یک دقیقه آماده می‌شوند. زمان پردازش با طول فایل افزایش می‌یابد — یک ضبط 10 دقیقه‌ای معمولاً در 20 تا 40 ثانیه و یک ضبط یک‌ساعته در حدود 4 تا 6 دقیقه نتیجه می‌دهد. فایل‌های طولانی‌تر به‌صورت خودکار تقسیم، به‌صورت موازی رونویسی، و سپس ادغام می‌شوند، پس نیازی نیست قبل از آپلود فایل صوتی‌تان را ببرید.
آیا می‌توانم صدای فارسی را به انگلیسی یا زبان‌های دیگر ترجمه کنم؟
بله، به دو روش. اگر فقط به خروجی انگلیسی نیاز دارید، گزینه ترجمه به انگلیسی را در تنظیمات آپلود فعال کنید — OpenAI Whisper صدای فارسی را در یک مرحله رونویسی و به متن انگلیسی ترجمه می‌کند، بدون هزینه اضافی. برای هر زبان مقصد دیگری، ابتدا فایل را به‌صورت معمول رونویسی کنید، سپس از گزینه ترجمه روی متن نهایی در داشبوردتان استفاده کنید — این قابلیت با Claude AI کار می‌کند و می‌تواند متن رونویسی‌شده فارسی (یا خلاصه و فهرست اقدامات آن) را به هر یک از زبان‌های پشتیبانی‌شده در convert.express ترجمه کند، نه فقط انگلیسی.
برای رونویسی فارسی از کدام مدل استفاده کنم؟
OpenAI Whisper و Gemini 3.5 Transcribe هر دو از زبان فارسی پشتیبانی می‌کنند. OpenAI Whisper گزینه‌ای شناخته‌شده است با پوشش گسترده زبانی و قابلیت ترجمه مستقیم به انگلیسی در یک مرحله. Gemini 3.5 Transcribe در مرحله پیش‌نمایش است، بیشترین پوشش زبانی را دارد و گویندگان را درون مدل برچسب‌گذاری می‌کند. MAI-Transcribe 1.5 از زبان فارسی پشتیبانی نمی‌کند. روی یک کلیپ کوتاه هر دو را امتحان کنید تا ببینید کدام خروجی برای ضبط شما مناسب‌تر است.
چه فرمت‌های صوتی و تصویری پشتیبانی می‌شوند؟
فرمت‌های MP3، WAV، M4A، FLAC، OGG، Opus، MP4، MOV، MKV و اکثر فرمت‌های رایج صوتی و تصویری پذیرفته می‌شوند. فایل‌هایی تا حجم 2 GB قابل آپلود هستند. زبان به‌صورت خودکار از روی صدا تشخیص داده می‌شود و نیازی به انتخاب دستی زبان نیست.