Z.aiGLM-5.3-Flash

External Link

GLM-5.3-Flash is a native multimodal model with 320B total and 18B active parameters, a 1M-token context window, and always-on reasoning.

Context
1,000,000 tokens
Inputs
Text, Image, Video, Document
Outputs
Text
Released
Sep 2, 2026

Providers & Pricing

Estimate workload cost →
ProviderInput / 1MOutput / 1MCache read / 1MCache write / 1MObservedEvidence
$0.15$0.50$0.030Sep 2, 2026Provider-reported
$0.075$0.25$0.015Aug 29, 2026Provider-reported
GLM-5.3-FlashZ.aiLiveBench2026-06-2573.27percent34,7071,270glm-5.3-flashaverage per caseRecomputedwinner eligibleThird-party benchmark
GLM-5.3-FlashZ.aiLMArena Text Arenatext-2026-09-01-0115087206961,470.91rating4,399glm-5.3-flash; 95% CI [1461.80526636, 1480.01283657]; votes 4399; rank 23unknownOriginal sourcewinner eligibleThird-party benchmark
GLM-5.3-FlashZ.aiLMArena Vision Arenavision-2026-08-27-0115087206961,296.37rating1,389glm-5.3-flash; 95% CI [1279.58251138, 1313.15015420]; votes 1389; rank 15unknownOriginal sourcewinner eligibleThird-party benchmark

At a Glance

Model Facts

Identity

Status
active
Developer ID
glm-5.3-flash
Released
Sep 2, 2026
Version
GLM-5.3-Flash
License
MIT

Capacity

Context
1,000,000 tokens
Maximum output
131,072 tokens
Total parameters
320,000,000,000
Active parameters
18,000,000,000
Knowledge cutoff
Not reported

Interface and Access

Inputs
Text, Image, Video, Document
Outputs
Text
API available
Yes
Open weights
Yes
Self-hostable
Yes
Reasoning
Yes
Vision
Yes
Tool calling
Yes
Capabilities
agents, chat, computer-use, reasoning, structured outputs, tools, vision

This page represents one immutable developer model ID. Serving endpoints, pricing, regions, quantizations, and service tiers attach separately and never create duplicate model records.

Version Lineage

PredecessorUnknown
SuccessorsUnknown
Aliasesglm-5.3-flash, glm-5.3-flash, ox-alpha, zai-org/GLM-5.3-Flash

Explicit Unknowns

  • Tokenizer
  • Knowledge cutoff

Primary Evidence

Sources and Observation Date

Comparable Peers

Related Model Comparisons

All comparisons →
APairBContext
vscross-developer peersimage, text
vsfamily variantsimage, text, video
vscross-developer peersimage, text, video
vscross-developer peersimage, text
vscross-developer peersimage, text
vsfamily variantsimage, text
vscross-developer peerstext
vscross-developer peerstext
vsfamily variantstext
vsfamily variantstext
vscross-developer peerstext
vscross-developer peerstext

Questions

GLM-5.3-Flash FAQs

What is GLM-5.3-Flash?+

GLM-5.3-Flash is a native multimodal model with 320B total and 18B active parameters, a 1M-token context window, and always-on reasoning. It is developed by Z.ai and its current sourced lifecycle status is active.

What inputs and outputs does GLM-5.3-Flash support?+

GLM-5.3-Flash's current record lists Text, Image, Video, and Document as input and Text as output.

How much context does GLM-5.3-Flash support?+

The current record lists a 1,000,000-token context window and a maximum output of 131,072 tokens.

Is GLM-5.3-Flash available through an API?+

GLM-5.3-Flash is marked as API-available. The directly sourced provider records currently include Deepinfra and Z.ai.

How large is GLM-5.3-Flash?+

The current record lists 320,000,000,000 total parameters and 18,000,000,000 active parameters. Its architecture is hybrid_sparse_linear_attention_moe.

What capabilities and tasks does GLM-5.3-Flash support?+

GLM-5.3-Flash's recorded capabilities are agents, chat, computer-use, reasoning, structured outputs, tools, and vision. Its supported tasks are coding, computer-use, document-question-answering, image, text, and video.

How much does GLM-5.3-Flash cost through an API?+

The lowest directly sourced prices currently attached to GLM-5.3-Flash are $0.075 per million input tokens through Z.ai and $0.25 per million output tokens through Z.ai. Prices are provider-specific and should be checked against each cited observation date.

Can GLM-5.3-Flash be self-hosted?+

GLM-5.3-Flash is an open-weight model. Self-hosting is supported, and the recorded license is MIT.

Does GLM-5.3-Flash have related versions or aliases?+

No predecessor is recorded. No successor is recorded. Recorded aliases are glm-5.3-flash, ox-alpha, and zai-org/GLM-5.3-Flash.

Which models can GLM-5.3-Flash be compared with?+

The current compatible peer set includes DeepSeek-V4-Flash-Vision-Exp, GLM-5V-Turbo, Nova 2 Lite, Mistral Large 3, and Command A Vision. Compatibility requires shared sourced modalities and tasks; it does not imply that one model is better. Browse model comparisons

Send Feedback