MAI-Transcribe-2: an hour of audio transcribed for 10 cents, across 60 languages
MAI-Transcribe-2 is Microsoft AI's in-house speech-to-text model, converting audio into text across 60 languages. Released on September 3, 2026 in public preview, it is priced at $0.10 per hour of audio, a launch rate that lasts until the end of 2026. At release it topped the multilingual FLEURS benchmark. One useful caveat, this is an API for developers rather than an app you install.
- Topped the FLEURS benchmark at launch across 60 languages
- Speaker diarization and word-level timestamps included by default
- Up to 10x faster on long recordings
- Launch price of $0.10 per audio hour
- Keyword biasing for industry jargon and proper names
- Still in public preview, no service-level agreement
- Built for developers, no ready-made consumer app
- Files capped at 300 MB or two hours
Speaker labels, word timestamps and your own jargon
MAI-Transcribe-2 separates the people talking in a recording, attributes each sentence to the right voice and stamps every word with precise timing. You upload a four-person steering call, you get back a labeled transcript ready for search, quotes or subtitles.
Language detection is automatic, and the model follows conversations that flip between two languages mid-sentence (Hinglish, the Hindi-English blend, appears in Microsoft's own demos). The company also pitches it as the listening layer for voice AI tools, paired with its MAI-Voice-2 Flash speech model on the output side.
- Keyword biasing, feed it product names and jargon to sharpen recognition
- Verbatim style, every filler word kept for compliance and analysis
- Clean style, readable text with the ums stripped out
MAI-Transcribe-2 against Whisper, GPT-Transcribe and Scribe v2
On the multilingual FLEURS benchmark, MAI-Transcribe-2 took first place at launch with a 5.2% average word error rate across 60 languages, ahead of Whisper-Large-V3, OpenAI's GPT-Transcribe, ElevenLabs' Scribe v2 and Google's Gemini 3.5 Transcribe. Independent tester Artificial Analysis also clocked it at up to ten times the speed of rivals on long recordings.
Three releases in five months, that is the pace of this family. Microsoft is building its own speech recognition stack in-house, step by step reducing its reliance on OpenAI's models.
| Version | Released | Languages | Highlights |
|---|---|---|---|
| MAI-Transcribe-1 | April 2026 | 25 | launched at $0.36 per hour |
| MAI-Transcribe-1.5 | June 2026 | 43 | keyword biasing, faster processing |
| MAI-Transcribe-2 | September 2026 | 60 | diarization, timestamps, output styles |
What MAI-Transcribe-2 costs and where to try it
MAI-Transcribe-2 costs $0.10 per hour of audio, a limited-time rate valid through December 31, 2026. The first generation launched at $0.36, so the cut reaches 72%. For context, Artificial Analysis lists Scribe v2 around $0.22 per hour, GPT-Transcribe near $0.27 and Gemini 3.5 Transcribe close to $0.30.
Access runs through Microsoft Foundry with an Azure account (free to create, no small print there), through the MAI Playground for code-free tests, or via OpenRouter. Since the January price remains unannounced, double-check the official page before locking in next year's budget.
Frequently asked questions
Is MAI-Transcribe-2 free?
No, it costs $0.10 per hour of audio, a promotional rate through the end of 2026. Creating an Azure account is free, and the MAI Playground lets you test the model without writing code. Microsoft has not yet published the standard rate that will follow.
MAI-Transcribe-2 vs Whisper, which should you pick?
Whisper remains the open-source option you can self-host, handy when audio must stay on your own machines. On the FLEURS benchmark, MAI-Transcribe-2 posts a lower word error rate than Whisper-Large-V3 and ships with diarization and timestamps built in, features Whisper does not provide on its own.
Can you use MAI-Transcribe-2 in production?
Microsoft labels it a public preview without a service-level agreement, so production workloads are officially discouraged for now. Worth knowing, MAI-Transcribe-1 was deprecated five months after release, so plan migration windows if you build on this model family.
What file formats and sizes does it accept?
WAV, MP3 and FLAC files, up to 300 MB or two hours each. Word-level timestamps and speaker labels are switched on through API parameters, and diarization in enhanced mode supports shorter recordings than plain transcription does.
Verdict: No consumer app here, just an API to wire in. Developers and data teams processing multilingual audio at scale get labeled transcripts for roughly a third of rival rates, and the MAI Playground makes the first test painless.
