MicrosoftVibeVoice-ASR-Streaming-1.5B

External Link

VibeVoice-ASR-Streaming-1.5B is an open-weight streaming automatic-speech-recognition model from Microsoft that transcribes speaker-attributed speech, accepts customized hotwords, and supports ten languages.

Context
Unknown
Inputs
Audio
Outputs
Text
Released
Sep 3, 2026

Providers & Pricing

Estimate workload cost →

At a Glance

Model Facts

Identity

Status
active
Developer ID
microsoft/VibeVoice-ASR-Streaming-1.5B
Released
Sep 3, 2026
Version
VibeVoice-ASR-Streaming-1.5B
License
mit

Capacity

Context
Unknown
Maximum output
Unknown
Total parameters
2,814,116,321
Active parameters
Unknown
Knowledge cutoff
Not reported

Interface and Access

Inputs
Audio
Outputs
Text
API available
No
Open weights
Yes
Self-hostable
Yes
Reasoning
No
Vision
No
Tool calling
No
Capabilities
diarization, hotwords, streaming, transcription

This page represents one immutable developer model ID. Serving endpoints, pricing, regions, quantizations, and service tiers attach separately and never create duplicate model records.

Version Lineage

PredecessorUnknown
SuccessorsUnknown
Aliasesmicrosoft/VibeVoice-ASR-Streaming-1.5B, vibevoice-asr-streaming-1.5b, VibeVoice ASR Streaming 1.5B

Explicit Unknowns

  • Maximum output
  • Tokenizer
  • Knowledge cutoff

Primary Evidence

Sources and Observation Date

Comparable Peers

Related Model Comparisons

All comparisons →
APairBContext
vsfamily variantsaudio, text
vsdeveloper peersaudio, text
vscross-developer peersaudio, text
vscross-developer peersaudio, text
vscross-developer peersaudio, text
vscross-developer peersaudio, text
vscross-developer peerstext
vscross-developer peerstext
vscross-developer peerstext
vscross-developer peersaudio
vscross-developer peersaudio

Questions

VibeVoice-ASR-Streaming-1.5B FAQs

What is VibeVoice-ASR-Streaming-1.5B?+

VibeVoice-ASR-Streaming-1.5B is an open-weight streaming automatic-speech-recognition model from Microsoft that transcribes speaker-attributed speech, accepts customized hotwords, and supports ten languages. It is developed by Microsoft and its current sourced lifecycle status is active.

What inputs and outputs does VibeVoice-ASR-Streaming-1.5B support?+

VibeVoice-ASR-Streaming-1.5B's current record lists Audio as input and Text as output.

Is VibeVoice-ASR-Streaming-1.5B available through an API?+

VibeVoice-ASR-Streaming-1.5B is explicitly marked as not API-available in the current record.

How much does VibeVoice-ASR-Streaming-1.5B cost through an API?+

No directly sourced token price is currently attached to VibeVoice-ASR-Streaming-1.5B. Missing prices remain unknown rather than being inferred from another model or provider.

How large is VibeVoice-ASR-Streaming-1.5B?+

The current record lists 2,814,116,321 total parameters and an unknown active parameter count. Its architecture is VibeVoiceForASRStreamingTraining.

What capabilities and tasks does VibeVoice-ASR-Streaming-1.5B support?+

VibeVoice-ASR-Streaming-1.5B's recorded capabilities are diarization, hotwords, streaming, and transcription. Its supported tasks are audio and transcription.

Can VibeVoice-ASR-Streaming-1.5B be self-hosted?+

VibeVoice-ASR-Streaming-1.5B is an open-weight model. Self-hosting is supported, and the recorded license is mit.

Does VibeVoice-ASR-Streaming-1.5B have related versions or aliases?+

No predecessor is recorded. No successor is recorded. Recorded aliases are microsoft/VibeVoice-ASR-Streaming-1.5B, vibevoice-asr-streaming-1.5b, and VibeVoice ASR Streaming 1.5B.

Which models can VibeVoice-ASR-Streaming-1.5B be compared with?+

The current compatible peer set includes VibeVoice-ASR-Streaming-7B, Phi-4-multimodal-instruct, Granite Speech 5.0 TurboCTC, Granite Speech 5.0 TurboCTC NC, and Inkling. Compatibility requires shared sourced modalities and tasks; it does not imply that one model is better. Browse model comparisons

How current is the Model Markets record for VibeVoice-ASR-Streaming-1.5B?+

VibeVoice-ASR-Streaming-1.5B was last verified Sep 2, 2026 from Microsoft's primary source, “microsoft/VibeVoice-ASR-Streaming-1.5B developer model repository.” The evidence is classified as official fact.

Send Feedback