Identity
- Status
- preview
- Developer ID
mai-transcribe- Released
- Sep 3, 2026
- Version
- MAI-Transcribe 2
- License
- Unknown
MAI-Transcribe 2 is Microsoft's preview speech-to-text model for multilingual transcription with speaker diarization and word-level timestamps.
| Provider | Price dimension | Amount | Unit | Observed | Evidence |
|---|---|---|---|---|---|
| Transcription | $0.10 | per hour | Sep 3, 2026 | Provider-reported |
At a Glance
mai-transcribeThis page represents one developer model product. Dated API snapshots, serving endpoints, pricing, regions, quantizations, and service tiers attach as versioned aliases or provider details and never create duplicate public model records.
Primary Evidence
Comparable Peers
| A | Pair | B | Context |
|---|---|---|---|
MAI-Transcribe 2Microsoft | vs | VibeVoice-ASR-Streaming-1.5BMicrosoft | developer peersaudio, text |
MAI-Transcribe 2Microsoft | vs | VibeVoice-ASR-Streaming-7BMicrosoft | developer peersaudio, text |
MAI-Transcribe 2Microsoft | vs | cross-developer peersaudio, text | |
MAI-Transcribe 2Microsoft | vs | cross-developer peersaudio, text | |
MAI-Transcribe 2Microsoft | vs | Lyria 3.5Google DeepMind | cross-developer peersaudio, text |
MAI-Transcribe 2Microsoft | vs | Gemini 2.5 FlashGoogle DeepMind | cross-developer peerstext |
MAI-Transcribe 2Microsoft | vs | Gemini 2.5 Flash-LiteGoogle DeepMind | cross-developer peerstext |
MAI-Transcribe 2Microsoft | vs | Gemini 2.5 ProGoogle DeepMind | cross-developer peerstext |
MAI-Transcribe 2Microsoft | vs | Gemini 3 FlashGoogle DeepMind | cross-developer peerstext |
MAI-Transcribe 2Microsoft | vs | MiniMax-Music3MiniMax | cross-developer peersaudio |
MAI-Transcribe 2Microsoft | vs | stable-audio-3-mediumStability AI | cross-developer peersaudio |
Questions
MAI-Transcribe 2 is Microsoft's preview speech-to-text model for multilingual transcription with speaker diarization and word-level timestamps. It is developed by Microsoft and its current sourced lifecycle status is preview.
MAI-Transcribe 2's current record lists Audio as input and Text as output.
MAI-Transcribe 2 is marked as API-available. The directly sourced provider records currently include Microsoft Foundry.
No directly sourced token price is currently attached to MAI-Transcribe 2. Missing prices remain unknown rather than being inferred from another model or provider.
MAI-Transcribe 2's recorded capabilities are diarization, multilingual, speaker-attribution, transcription, and word-level-timestamps. Its supported tasks are audio and transcription.
MAI-Transcribe 2 is not marked as an open-weight model. Self-hosting is not established by the current source, with no license recorded.
No predecessor is recorded. No successor is recorded. Recorded aliases are mai-transcribe, mai-transcribe-2, and MAI-Transcribe-2.
The current compatible peer set includes VibeVoice-ASR-Streaming-1.5B, VibeVoice-ASR-Streaming-7B, Granite Speech 5.0 TurboCTC, Granite Speech 5.0 TurboCTC NC, and Lyria 3.5. Compatibility requires shared sourced modalities and tasks; it does not imply that one model is better. Browse model comparisons →
MAI-Transcribe 2 was last verified Sep 3, 2026 from Microsoft's primary source, “MAI-Transcribe-2: Highest quality transcription, at the fastest speed and lowest cost.” The evidence is classified as official fact.