VibeVoice-ASR-Streaming-1.5B vs MAI-Transcribe 2

Why this pair: Microsoft hosted and open-weight speech-recognition models

Benchmark Performance

Available Benchmarks

No Protocol-Matched Benchmark Yet.Results appear here only when both models share the same benchmark version, metric, evaluation protocol, and evidence class.
FieldAt a Glance
Microsoft · activeVibeVoice-ASR-Streaming-1.5BVerified Sep 2, 2026
Microsoft · previewMAI-Transcribe 2Verified Sep 3, 2026

Technical Differences

Side-by-Side Facts

CuratedIndexable
FieldVibeVoice-ASR-Streaming-1.5BMAI-Transcribe 2
DeveloperMicrosoftMicrosoft
FamilyVibevoice ASR StreamingMai Transcribe
ModelVibeVoice-ASR-Streaming-1.5BMAI-Transcribe 2
VersionVibeVoice-ASR-Streaming-1.5BMAI-Transcribe 2
Lifecycleactivepreview
Released2026-09-032026-09-03
Knowledge cutoffUnknownUnknown
Input modalitiesAudioAudio
Output modalitiesTextText
Context windowUnknownUnknown
Total parameters2,814,116,321Unknown
Active parametersUnknownUnknown
LicensemitUnknown
Open weightsYesNo
API availableNoYes
Self-hostableYesUnknown
Provider accessUnknownMicrosoft Foundry (Public preview)
Capabilitiesdiarization, hotwords, multilingual, streaming, transcriptiondiarization, multilingual, speaker-attribution, transcription, word-level-timestamps
Streaming chunk22 framesUnknown
Streaming lookahead4 framesUnknown
Supported languagesChinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, SpanishUnknown
Target sample rate24000 HzUnknown

11 comparable fields · 7 material differences · Editorially curated pair passes the primary-source comparison gate

VibeVoice-ASR-Streaming-1.5B Capabilities

diarizationhotwordsmultilingualstreamingtranscription
Input price
Output price
Serving providers0
Canonical IDmicrosoft/VibeVoice-ASR-Streaming-1.5B

MAI-Transcribe 2 Capabilities

diarizationmultilingualspeaker-attributiontranscriptionword-level-timestamps
Input price
Output price
Serving providers1
Canonical IDmicrosoft/mai-transcribe

Internal Comparison Graph

Related Comparisons

All audio comparisons →
APairBContext
vsfamily variantsaudio, text
vsdeveloper peersaudio, text
vscross-developer peersaudio, text
vsdeveloper peersaudio, text
vscross-developer peersaudio, text
vscross-developer peersaudio, text
vscross-developer peersaudio, text
vscross-developer peersaudio, text
vscross-developer peersaudio, text
vscross-developer peersaudio, text
vscross-developer peersaudio, text
vscross-developer peersaudio

Primary Evidence

Sources and Freshness

Questions

VibeVoice-ASR-Streaming-1.5B vs MAI-Transcribe 2 FAQs

Is VibeVoice-ASR-Streaming-1.5B or MAI-Transcribe 2 better for coding?+

This comparison does not currently contain a protocol-matched coding benchmark for both VibeVoice-ASR-Streaming-1.5B and MAI-Transcribe 2, so Model Markets cannot name a coding leader from pricing, context size, or capability labels alone.

Which is cheaper, VibeVoice-ASR-Streaming-1.5B or MAI-Transcribe 2?+

Neither model has a directly sourced input price in this comparison. Neither model has a directly sourced output price in this comparison.

Which has a larger context window, VibeVoice-ASR-Streaming-1.5B or MAI-Transcribe 2?+

Neither model has a larger sourced context window in this comparison. VibeVoice-ASR-Streaming-1.5B is — and MAI-Transcribe 2 is —.

Which performs better in benchmarks, VibeVoice-ASR-Streaming-1.5B or MAI-Transcribe 2?+

There is no overall benchmark winner: At least two independently verified, protocol-matched benchmarks are required for an overall winner.

Can VibeVoice-ASR-Streaming-1.5B or MAI-Transcribe 2 be self-hosted?+

VibeVoice-ASR-Streaming-1.5B is the only model in this pair currently marked as self-hostable. VibeVoice-ASR-Streaming-1.5B is open weight; MAI-Transcribe 2 is not marked open weight.

Can VibeVoice-ASR-Streaming-1.5B and MAI-Transcribe 2 understand images?+

VibeVoice-ASR-Streaming-1.5B is not documented with image input; MAI-Transcribe 2 is not documented with image input. This reflects supported input modalities, not vision quality.

Which can generate longer answers, VibeVoice-ASR-Streaming-1.5B or MAI-Transcribe 2?+

Neither has a larger sourced maximum output. VibeVoice-ASR-Streaming-1.5B is — and MAI-Transcribe 2 is —.

Do VibeVoice-ASR-Streaming-1.5B and MAI-Transcribe 2 support reasoning and tool use?+

VibeVoice-ASR-Streaming-1.5B: none of these features are definitively sourced. MAI-Transcribe 2: none of these features are definitively sourced. Feature support does not establish relative quality.

Which is available from more inference providers, VibeVoice-ASR-Streaming-1.5B or MAI-Transcribe 2?+

VibeVoice-ASR-Streaming-1.5B has 0 sourced provider routes; MAI-Transcribe 2 has 1, so MAI-Transcribe 2 has broader tracked availability.

Which offers better value, VibeVoice-ASR-Streaming-1.5B or MAI-Transcribe 2?+

There is no universal value winner. Compare the input and output prices above with the matched benchmark result for your workload: cheaper tokens can be offset by different quality, token usage, latency, or provider availability.

Send Feedback