QwenQwen3.8-Flash

External Link

Qwen3.8-Flash is a low-latency multimodal model with a 1M-token context window and OpenAI- and Anthropic-compatible APIs.

Context
1,000,000 tokens
Inputs
Text, Image, Video
Outputs
Text
Released
Not reported

Providers & Pricing

Estimate workload cost →
ProviderInput / 1MOutput / 1MCache read / 1MCache write / 1MObservedEvidence
$0.80$2.70$0.10Aug 29, 2026Provider-reported
$0.15$0.47$0.016Sep 2, 2026Provider-reported

At a Glance

Model Facts

Identity

Status
active
Developer ID
qwen3.8-flash
Released
Not reported
Version
Qwen3.8-Flash
License
Unknown

Capacity

Context
1,000,000 tokens
Maximum output
131,072 tokens
Total parameters
Unknown
Active parameters
Unknown
Knowledge cutoff
Not reported

Interface and Access

Inputs
Text, Image, Video
Outputs
Text
API available
Yes
Open weights
No
Self-hostable
No
Reasoning
Yes
Vision
Yes
Tool calling
Yes
Capabilities
agents, chat, computer-use, reasoning, structured outputs, tools, vision

This page represents one immutable developer model ID. Serving endpoints, pricing, regions, quantizations, and service tiers attach separately and never create duplicate model records.

Version Lineage

PredecessorUnknown
SuccessorsUnknown
Aliasesqwen3.8-flash

Explicit Unknowns

  • Release date
  • Architecture
  • Tokenizer
  • Knowledge cutoff

Primary Evidence

Sources and Observation Date

Comparable Peers

Related Model Comparisons

All comparisons →
APairBContext
vsfamily variantsimage, text, video
vsfamily variantsimage, text, video
vscross-developer peersimage, text, video
vscross-developer peerstext
vscross-developer peerstext
vscross-developer peerstext
vscross-developer peerstext
vscross-developer peerstext
vscross-developer peerstext
vscross-developer peerstext
vsfamily variantstext
vsfamily variantstext

Questions

Qwen3.8-Flash FAQs

What is Qwen3.8-Flash?+

Qwen3.8-Flash is a low-latency multimodal model with a 1M-token context window and OpenAI- and Anthropic-compatible APIs. It is developed by Qwen and its current sourced lifecycle status is active.

What inputs and outputs does Qwen3.8-Flash support?+

Qwen3.8-Flash's current record lists Text, Image, and Video as input and Text as output.

How much context does Qwen3.8-Flash support?+

The current record lists a 1,000,000-token context window and a maximum output of 131,072 tokens.

Is Qwen3.8-Flash available through an API?+

Qwen3.8-Flash is marked as API-available. The directly sourced provider records currently include Alibaba Cloud Model Studio and Openrouter.

How much does Qwen3.8-Flash cost through an API?+

The lowest directly sourced prices currently attached to Qwen3.8-Flash are $0.15 per million input tokens through Openrouter and $0.47 per million output tokens through Openrouter. Prices are provider-specific and should be checked against each cited observation date.

What capabilities and tasks does Qwen3.8-Flash support?+

Qwen3.8-Flash's recorded capabilities are agents, chat, computer-use, reasoning, structured outputs, tools, and vision. Its supported tasks are coding, computer-use, image, reasoning, text, and video.

Can Qwen3.8-Flash be self-hosted?+

Qwen3.8-Flash is not marked as an open-weight model. Self-hosting is marked as unsupported, with no license recorded.

Does Qwen3.8-Flash have related versions or aliases?+

No predecessor is recorded. No successor is recorded. Recorded aliases are qwen3.8-flash.

Which models can Qwen3.8-Flash be compared with?+

The current compatible peer set includes Qwen3.7-Plus, Qwen3.8-Max, GLM-5V-Turbo, Gemini 2.5 Flash, and Gemini 3.5 Flash. Compatibility requires shared sourced modalities and tasks; it does not imply that one model is better. Browse model comparisons

How current is the Model Markets record for Qwen3.8-Flash?+

Qwen3.8-Flash was last verified Aug 29, 2026 from Qwen's primary source, “qwen3.8-flash Model Info | Alibaba Cloud Model Studio.” The evidence is classified as official fact.

Send Feedback