Transcribe Persian Audio or Video to Text
Upload any Persian recording — OpenAI Whisper detects the language automatically and returns an accurate transcript in seconds.
Drag a file here or browse
MP3, WAV, M4A, FLAC · MP4, MOV, MKV, WebM · up to 2 GB
Set language, translation, or vocabulary hints below before transcribing
YouTube, Dropbox, Google Drive, or any direct MP3/MP4 link
Not applied on Gemini 3.5 Transcribe — the model can't combine vocabulary hints with word-level timestamps
MP3, WAV, M4A, FLAC · MP4, MOV, MKV · up to 2 GB · 10 minutes free every day
About Persian transcription
Persian sits within the Iranian branch of the Indo-European family, which gives it a relatively regular phonological structure compared to many of its geographic neighbors. That regularity helps: spoken Farsi has consistent consonant-vowel patterns and no lexical tone, so a speech recognition system doesn't have to resolve pitch ambiguity the way it would for a tonal language. The Persian alphabet, a variant of Arabic script, omits short vowels in standard writing, but that's a challenge for text rendering rather than for audio-to-text conversion, where the acoustic signal carries the vowel information directly. The upshot is that standard spoken Iranian Farsi transcribes with high accuracy. The one area worth knowing about is the gap between Iranian Farsi and Afghan Dari. They share virtually all their grammar, but vocabulary and pronunciation diverge enough that a system trained narrowly on Tehran speech would stumble on Kabul speech. Convert.express covers both varieties, so researchers, journalists, and legal teams working with Dari recordings don't need a separate workflow.
OpenAI Whisper and Gemini 3.5 Transcribe both support Persian, detected automatically from the audio — no manual language selection is needed. MAI-Transcribe 1.5 does not support this language. Upload any Persian recording up to 2 GB — MP3, WAV, M4A, FLAC, MP4, MOV, and most common audio and video formats are accepted. Every account includes 10 free minutes of transcription per day; longer files are billed at €0.02 per minute with a €0.50 minimum per transaction. Transcripts download as plain text and are automatically deleted within 24 hours of upload.
Available transcription models
10 min
Free every day
€0.02
Per minute beyond quota
< 1 min
Typical turnaround
Tips for the best results
- Use a quiet environment — background noise is the biggest source of transcription errors.
- Speak at a natural pace; extremely fast speech reduces accuracy in any language.
- MP3 or M4A files work great; uncompressed WAV is ideal for professional recordings.
- Files up to 2 GB are supported — long recordings are automatically split and merged.
Frequently asked questions
- How accurate is AI transcription for Persian?
- Accuracy for Persian is high on convert.express, and the linguistic profile explains why. Iranian Farsi has no lexical tone and a regular sound system, which removes two common sources of transcription error. The Persian alphabet's omission of short vowels is a writing convention that doesn't affect audio processing. The bigger variable is dialect: Iranian Farsi and Afghan Dari diverge in vocabulary and pronunciation despite sharing the same grammar, and the service covers both, so you can expect consistent results whether your audio comes from Tehran or Kabul.
- How long does Persian transcription take?
- Most Persian recordings are ready in under a minute. Processing time scales with file length — a 10-minute recording typically returns results in 20–40 seconds, and a one-hour recording in around 4–6 minutes. Longer files are automatically split, transcribed in parallel, and merged, so you never need to cut your audio before uploading.
- Can I translate Persian audio to English or other languages?
- Yes, two ways. For a quick English-only result, enable "Translate to English" in the upload options above — OpenAI Whisper transcribes and translates Persian audio to English text in a single pass, at no extra cost. For any other target language, first transcribe normally, then use the Translate action on the finished transcript in your dashboard — it's powered by Claude AI and can translate the Persian transcript (or its summary and action list) into any of convert.express's other supported languages, not just English.
- Which transcription model should I use for Persian?
- OpenAI Whisper and Gemini 3.5 Transcribe both support Persian. OpenAI Whisper is the established choice, with broad language coverage and single-pass translation to English. Gemini 3.5 Transcribe is in preview, has the widest language coverage, and labels speakers in-model. MAI-Transcribe 1.5 does not support Persian. Try them on a short clip to see which output suits your recording.
- What audio and video formats are supported?
- MP3, WAV, M4A, FLAC, OGG, Opus, MP4, MOV, MKV, and most other common audio and video formats are accepted. Files up to 2 GB can be uploaded. The language is detected automatically from the audio — no manual language selection is needed.
فارسی · convert.express
زبان فارسی شاخهای از خانواده هند-اروپایی است و در مقایسه با بسیاری از زبانهای همسایه، ساختار آوایی نسبتاً منظمی دارد. این ویژگی به نفع تشخیص گفتار است: فارسی محاورهای الگوهای همخوان-واکهای یکدستی دارد و فاقد تکیه واژگانی است، بنابراین سیستم تشخیص گفتار نیازی به رفع ابهام زیروبمی ندارد — برخلاف زبانهای تکیهدار که این کار را پیچیده میکنند. الفبای فارسی، که گونهای از خط عربی است، در نوشتار استاندارد مصوتهای کوتاه را حذف میکند؛ اما این موضوع چالش رندرینگ متن است، نه تبدیل صدا به نوشتار — چون در پردازش صوتی، اطلاعات مصوت مستقیماً از سیگنال صوتی به دست میآید. نتیجه این است که فارسی گفتاری استاندارد ایرانی با دقت بالایی رونویسی میشود. یک نکته مهم که باید بدانید تفاوت میان فارسی ایرانی و دری افغانستانی است. این دو زبان تقریباً دستور زبان مشترکی دارند، اما واژگان و تلفظشان به اندازهای با هم متفاوت است که سیستمی که صرفاً بر اساس گفتار تهرانی آموزش دیده، در برابر گفتار کابلی دچار مشکل میشود. convert.express هر دو گونه را پشتیبانی میکند، بنابراین محققان، روزنامهنگاران و تیمهای حقوقی که با ضبطهای دری کار میکنند نیازی به فرآیند جداگانهای ندارند. OpenAI Whisper و Gemini 3.5 Transcribe هر دو از زبان فارسی پشتیبانی میکنند و زبان بهصورت خودکار از روی صدا تشخیص داده میشود — نیازی به انتخاب دستی زبان نیست. MAI-Transcribe 1.5 این زبان را پشتیبانی نمیکند. هر فایل صوتی فارسی تا حجم 2 GB را میتوانید آپلود کنید — فرمتهای MP3، WAV، M4A، FLAC، MP4، MOV و اکثر فرمتهای رایج صوتی و تصویری پذیرفته میشوند. هر حساب کاربری روزانه 10 دقیقه رونویسی رایگان دارد؛ برای فایلهای طولانیتر، نرخ €0.02 به ازای هر دقیقه اعمال میشود و حداقل مبلغ هر تراکنش €0.50 است. فایلهای متنی رونویسیشده قابل دانلود هستند و ظرف 24 ساعت پس از آپلود بهصورت خودکار حذف میشوند.
- دقت رونویسی هوش مصنوعی برای زبان فارسی چقدر است؟
- دقت رونویسی فارسی در convert.express بالاست و ویژگیهای زبانی این زبان دلیل آن را روشن میکنند. فارسی ایرانی نه تکیه واژگانی دارد و نه دستگاه آوایی پیچیدهای، که همین دو عامل از رایجترین منابع خطا در رونویسی هستند. حذف مصوتهای کوتاه در الفبای فارسی یک قرارداد نوشتاری است و تأثیری بر پردازش صوتی ندارد. متغیر مهمتر گویش است: فارسی ایرانی و دری افغانستانی با وجود دستور زبان مشترک، در واژگان و تلفظ تفاوتهایی دارند، و این سرویس هر دو را پوشش میدهد — بنابراین چه صدایتان از تهران باشد چه از کابل، میتوانید انتظار نتایج یکسانی داشته باشید.
- رونویسی فارسی چقدر طول میکشد؟
- اکثر ضبطهای فارسی در کمتر از یک دقیقه آماده میشوند. زمان پردازش با طول فایل افزایش مییابد — یک ضبط 10 دقیقهای معمولاً در 20 تا 40 ثانیه و یک ضبط یکساعته در حدود 4 تا 6 دقیقه نتیجه میدهد. فایلهای طولانیتر بهصورت خودکار تقسیم، بهصورت موازی رونویسی، و سپس ادغام میشوند، پس نیازی نیست قبل از آپلود فایل صوتیتان را ببرید.
- آیا میتوانم صدای فارسی را به انگلیسی یا زبانهای دیگر ترجمه کنم؟
- بله، به دو روش. اگر فقط به خروجی انگلیسی نیاز دارید، گزینه ترجمه به انگلیسی را در تنظیمات آپلود فعال کنید — OpenAI Whisper صدای فارسی را در یک مرحله رونویسی و به متن انگلیسی ترجمه میکند، بدون هزینه اضافی. برای هر زبان مقصد دیگری، ابتدا فایل را بهصورت معمول رونویسی کنید، سپس از گزینه ترجمه روی متن نهایی در داشبوردتان استفاده کنید — این قابلیت با Claude AI کار میکند و میتواند متن رونویسیشده فارسی (یا خلاصه و فهرست اقدامات آن) را به هر یک از زبانهای پشتیبانیشده در convert.express ترجمه کند، نه فقط انگلیسی.
- برای رونویسی فارسی از کدام مدل استفاده کنم؟
- OpenAI Whisper و Gemini 3.5 Transcribe هر دو از زبان فارسی پشتیبانی میکنند. OpenAI Whisper گزینهای شناختهشده است با پوشش گسترده زبانی و قابلیت ترجمه مستقیم به انگلیسی در یک مرحله. Gemini 3.5 Transcribe در مرحله پیشنمایش است، بیشترین پوشش زبانی را دارد و گویندگان را درون مدل برچسبگذاری میکند. MAI-Transcribe 1.5 از زبان فارسی پشتیبانی نمیکند. روی یک کلیپ کوتاه هر دو را امتحان کنید تا ببینید کدام خروجی برای ضبط شما مناسبتر است.
- چه فرمتهای صوتی و تصویری پشتیبانی میشوند؟
- فرمتهای MP3، WAV، M4A، FLAC، OGG، Opus، MP4، MOV، MKV و اکثر فرمتهای رایج صوتی و تصویری پذیرفته میشوند. فایلهایی تا حجم 2 GB قابل آپلود هستند. زبان بهصورت خودکار از روی صدا تشخیص داده میشود و نیازی به انتخاب دستی زبان نیست.