Identity
- Status
- active
- Developer ID
microsoft/VibeVoice-ASR-Streaming-1.5B- Released
- Sep 3, 2026
- Version
- VibeVoice-ASR-Streaming-1.5B
- License
- mit
VibeVoice-ASR-Streaming-1.5B is an open-weight streaming automatic-speech-recognition model from Microsoft that transcribes speaker-attributed speech, accepts customized hotwords, and supports ten languages.
At a Glance
microsoft/VibeVoice-ASR-Streaming-1.5BThis page represents one immutable developer model ID. Serving endpoints, pricing, regions, quantizations, and service tiers attach separately and never create duplicate model records.
Primary Evidence
Comparable Peers
| A | Pair | B | Context |
|---|---|---|---|
VibeVoice-ASR-Streaming-1.5BMicrosoft | vs | VibeVoice-ASR-Streaming-7BMicrosoft | family variantsaudio, text |
VibeVoice-ASR-Streaming-1.5BMicrosoft | vs | Phi-4-multimodal-instructMicrosoft | developer peersaudio, text |
VibeVoice-ASR-Streaming-1.5BMicrosoft | vs | cross-developer peersaudio, text | |
VibeVoice-ASR-Streaming-1.5BMicrosoft | vs | cross-developer peersaudio, text | |
VibeVoice-ASR-Streaming-1.5BMicrosoft | vs | InklingThinking Machines Lab | cross-developer peersaudio, text |
VibeVoice-ASR-Streaming-1.5BMicrosoft | vs | Muse Spark 1.2Meta | cross-developer peersaudio, text |
VibeVoice-ASR-Streaming-1.5BMicrosoft | vs | Gemini 2.5 FlashGoogle DeepMind | cross-developer peerstext |
VibeVoice-ASR-Streaming-1.5BMicrosoft | vs | Gemini 2.5 Flash-LiteGoogle DeepMind | cross-developer peerstext |
VibeVoice-ASR-Streaming-1.5BMicrosoft | vs | Gemini 2.5 ProGoogle DeepMind | cross-developer peerstext |
VibeVoice-ASR-Streaming-1.5BMicrosoft | vs | MiniMax-Music3MiniMax | cross-developer peersaudio |
VibeVoice-ASR-Streaming-1.5BMicrosoft | vs | stable-audio-3-mediumStability AI | cross-developer peersaudio |
Questions
VibeVoice-ASR-Streaming-1.5B is an open-weight streaming automatic-speech-recognition model from Microsoft that transcribes speaker-attributed speech, accepts customized hotwords, and supports ten languages. It is developed by Microsoft and its current sourced lifecycle status is active.
VibeVoice-ASR-Streaming-1.5B's current record lists Audio as input and Text as output.
VibeVoice-ASR-Streaming-1.5B is explicitly marked as not API-available in the current record.
No directly sourced token price is currently attached to VibeVoice-ASR-Streaming-1.5B. Missing prices remain unknown rather than being inferred from another model or provider.
The current record lists 2,814,116,321 total parameters and an unknown active parameter count. Its architecture is VibeVoiceForASRStreamingTraining.
VibeVoice-ASR-Streaming-1.5B's recorded capabilities are diarization, hotwords, streaming, and transcription. Its supported tasks are audio and transcription.
VibeVoice-ASR-Streaming-1.5B is an open-weight model. Self-hosting is supported, and the recorded license is mit.
No predecessor is recorded. No successor is recorded. Recorded aliases are microsoft/VibeVoice-ASR-Streaming-1.5B, vibevoice-asr-streaming-1.5b, and VibeVoice ASR Streaming 1.5B.
The current compatible peer set includes VibeVoice-ASR-Streaming-7B, Phi-4-multimodal-instruct, Granite Speech 5.0 TurboCTC, Granite Speech 5.0 TurboCTC NC, and Inkling. Compatibility requires shared sourced modalities and tasks; it does not imply that one model is better. Browse model comparisons →
VibeVoice-ASR-Streaming-1.5B was last verified Sep 2, 2026 from Microsoft's primary source, “microsoft/VibeVoice-ASR-Streaming-1.5B developer model repository.” The evidence is classified as official fact.