PaddleOCR VL 1.5 vs VibeVoice ASR Streaming 1.5B
At a Glance
| Compare | PaddleOCR VL 1.5Baidu | VibeVoice ASR Streaming 1.5BMicrosoft |
|---|---|---|
| Pricing and Limits | ||
| Context windowMaximum documented tokens | 131K | Not reported |
| Model facts checked | Aug 28, 2026View model evidence → | Sep 2, 2026View model evidence → |
Token prices are the lowest available sourced USD rates; input and output may use different providers. Cost ranking estimates output spend on LiveBench, not a full request bill. Ranking methodology →
Available Benchmarks
Side-by-Side Facts
| Field | PaddleOCR-VL-1.5 | VibeVoice-ASR-Streaming-1.5B |
|---|---|---|
| Developer | Baidu | Microsoft |
| Family | Paddleocr VL 1 5 | Vibevoice ASR Streaming |
| Model | PaddleOCR-VL-1.5 | VibeVoice-ASR-Streaming-1.5B |
| Version | PaddleOCR-VL-1.5 | VibeVoice-ASR-Streaming-1.5B |
| Lifecycle | active | active |
| Released | 2026-01-29 | 2026-09-03 |
| Knowledge cutoff | Unknown | Unknown |
| Input modalities | Text, Image | Audio |
| Output modalities | Text | Text |
| Context window | 131K | Unknown |
| Total parameters | 958.6M | 2.8B |
| Active parameters | Unknown | Unknown |
| License | apache-2.0 | mit |
| Open weights | Yes | Yes |
| API available | Unknown | No |
| Self-hostable | Yes | Yes |
| Provider access | Unknown | Unknown |
| Capabilities | chat, generation | diarization, hotwords, multilingual, streaming, transcription |
| Streaming chunk | Unknown | 22 frames |
| Streaming lookahead | Unknown | 4 frames |
| Supported languages | Unknown | Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish |
| Target sample rate | Unknown | 24000 Hz |
PaddleOCR VL 1.5 Capabilities
VibeVoice ASR Streaming 1.5B Capabilities
Primary Evidence
Sources and Freshness
Questions
PaddleOCR VL 1.5 vs VibeVoice ASR Streaming 1.5B FAQs
Is PaddleOCR VL 1.5 or VibeVoice ASR Streaming 1.5B better for coding?+
This comparison does not currently contain a protocol-matched coding benchmark for both PaddleOCR VL 1.5 and VibeVoice ASR Streaming 1.5B, so Model Markets cannot name a coding leader from pricing, context size, or capability labels alone.
Which is cheaper, PaddleOCR VL 1.5 or VibeVoice ASR Streaming 1.5B?+
Neither model has a directly sourced input price in this comparison. Neither model has a directly sourced output price in this comparison.
Which has a larger context window, PaddleOCR VL 1.5 or VibeVoice ASR Streaming 1.5B?+
Neither model has a larger sourced context window in this comparison. PaddleOCR VL 1.5 is 131K and VibeVoice ASR Streaming 1.5B is —.
Which performs better in benchmarks, PaddleOCR VL 1.5 or VibeVoice ASR Streaming 1.5B?+
There is no overall benchmark winner: At least two independently verified, protocol-matched benchmarks are required for an overall winner.
Can PaddleOCR VL 1.5 or VibeVoice ASR Streaming 1.5B be self-hosted?+
Both models have the same recorded self-hosting status: supported. PaddleOCR VL 1.5 is open weight; VibeVoice ASR Streaming 1.5B is open weight.
Can PaddleOCR VL 1.5 and VibeVoice ASR Streaming 1.5B understand images?+
PaddleOCR VL 1.5 is documented with image input; VibeVoice ASR Streaming 1.5B is not documented with image input. This reflects supported input modalities, not vision quality.
Which can generate longer answers, PaddleOCR VL 1.5 or VibeVoice ASR Streaming 1.5B?+
Neither has a larger sourced maximum output. PaddleOCR VL 1.5 is — and VibeVoice ASR Streaming 1.5B is —.
Do PaddleOCR VL 1.5 and VibeVoice ASR Streaming 1.5B support reasoning and tool use?+
PaddleOCR VL 1.5: image input. VibeVoice ASR Streaming 1.5B: none of these features are definitively sourced. Feature support does not establish relative quality.
Which is available from more inference providers, PaddleOCR VL 1.5 or VibeVoice ASR Streaming 1.5B?+
PaddleOCR VL 1.5 has 0 sourced provider routes; VibeVoice ASR Streaming 1.5B has 0, a tie.
Which offers better value, PaddleOCR VL 1.5 or VibeVoice ASR Streaming 1.5B?+
There is no universal value winner. Compare the input and output prices above with the matched benchmark result for your workload: cheaper tokens can be offset by different quality, token usage, latency, or provider availability.