What Is GLM-5.3-Flash? Specs and Price
GLM-5.3-Flash is Z.ai's first natively multimodal model: 320B total and 18B active, MIT weights, and a ninth of GLM-5.3's price. Specs, scores, costs.
Table of Contents
GLM-5.3-Flash is Z.ai's first natively multimodal model, released 26 August 2026 with MIT-licensed weights, 320 billion total parameters and 18 billion active. It lists at $0.15 per million input tokens against flagship GLM-5.3's $1.40, and it spent the six days before launch running anonymously on OpenRouter as a stealth model called Ox Alpha.
Key Takeaways:
- Not a cut-down GLM-5.3. Separate architecture, separately trained base, and it accepts images where the flagship is text-only
- 320B total parameters with only 18B active, and 45 layers where GLM-4.5 had 92
- Weights are published under MIT. The flagship GLM-5.3's weights, promised for the same week, still are not
- It pulled 3.94 trillion prompt tokens during its anonymous preview, all served on Chinese AI chips
- Z.ai reports it approaching Claude Opus 4.8 on coding at roughly a ninth of the flagship's list price
The number Z.ai leads with is a cost-per-task figure: 57 on the Artificial Analysis Intelligence Index v4.1.1 at $0.045 per task on discounted pricing, a level the company says previously cost around ten times as much. Vendor framing, vendor benchmark selection, and still the most concrete claim in the launch: this is a model priced like a small one that does not benchmark like one.
GLM-5.3-Flash is not a smaller GLM-5.3
The naming invites exactly the wrong assumption. Flash releases are usually distillations of a flagship, so a reasonable reader sees "GLM-5.3-Flash" and assumes a compressed GLM-5.3. It isn't one.
| GLM-5.3 | GLM-5.3-Flash | |
|---|---|---|
| Base model | shared with GLM-5.2 | newly trained |
| Where the gains came from | post-training only | new architecture and corpus |
| Inputs | text only | text and images |
| Total / active parameters | not published | 320B / 18B |
| Weights | unpublished | MIT, on Hugging Face |
| List price in / out per 1M | $1.40 / $4.40 | $0.15 / $0.50 |
GLM-5.3 launched on 14 August reusing GLM-5.2's base model, with every improvement coming from post-training. GLM-5.3-Flash starts somewhere else entirely, with a redesigned architecture and a 30 trillion token multimodal pre-training corpus. Sharing a version number is a product decision.

Which one is better depends on what you are doing, and the answer inverts what the names imply for everything except peak coding. On Z.ai Code Bench at max effort, the flagship reaches 34.5% and Flash reaches 29.0%. On agentic benchmarks Flash is the stronger of the two in several places, and it is the only one of the pair that can look at a screenshot.
The architecture that makes it cheap
Three changes, and they compound.
Hybrid attention. Linear attention captures local dependencies through state modelling while sparse attention retrieves global context through a lightweight indexer. Z.ai adds IndexPool on top, which compresses four indexer key vectors into one through weighted pooling, specifically to hold down latency and memory at a million tokens of context. Against flagship GLM-5.3, the company reports attention compute down by a factor of 3.0 and KV cache size down by 4.4.
Fewer active parameters and far fewer layers. Compared with the GLM-4.5 series at a similar total size, 320B against 355B, Flash nearly halves both the activated parameters, 18B against 32B, and the layer count, 45 against 92.
Manifold-Constrained Hyper-Connections, which Z.ai credits for scaling efficiency without expanding on the mechanism in the launch post.
Z.ai is candid about where this leaves them: the KV cache is still slightly larger than Kimi-K3's and DeepSeek-V4-Flash's, and they say so in the same paragraph where they claim the lowest attention compute of the group. A launch post that names the metric it loses on is doing something unusual.
What GLM-5.3-Flash scores
The comparisons Z.ai puts in prose, which are safer to quote than its tables:
| Benchmark | GLM-5.3-Flash | GLM-5.2 |
|---|---|---|
| DeepSWE v1.1 | 63.4 | 46.2 |
| AutomationBench v1.0.6 | 48.8 | 26.2 |
| Toolathlon Verified | 78.4 | 59.9 |
| GDPval-AA v2 | 1773 | 1504 |
Against closed models the framing is "approaching" rather than beating, and Z.ai's own in-house evaluation puts a number on the gap: on Z.ai Code Bench v1.0 at max effort, run through Claude Code 2.1.207, GLM-5.3-Flash reaches 29.0 against Claude Opus 4.8's 29.5. Across the six benchmarks Z.ai charts, the split against Opus 4.8 is three and three: ahead on DeepSWE, AutomationBench and GDPval-AA, behind on Terminal Bench, Agents' Last Exam and HLE with tools.

The base model results are the ones I would watch. GLM-5.3-Flash-Base scores 37.6 on LiveCodeBench-Base against GLM-5-Base's 34.4 and DeepSeek-V4-Flash-Base's 29.9, while activating 18B parameters to GLM-5-Base's 40B. Post-training can be tuned toward a benchmark; a base model beating a larger one on code is a statement about the corpus and the architecture.
One caveat worth carrying. Z.ai's GLM-5.3 launch post two weeks earlier contained a comparison table whose column labels contradicted its own prose, so I am quoting the sentences rather than the grids from both.
Vision exists here for coding, not for image captioning
The multimodal part has a specific justification, and it is more interesting than "the model can see pictures now".
Z.ai's argument is that for frontend work, game development and 3D simulation, the deliverable is not code but a rendered interface, and many failures only surface through rendering or interaction. So the model needs to decide when to look at its own output. Their data pipelines target self-visual judgment and test-time improvement: trajectories where the model builds something, inspects the result, and refines it. For frontend coding they added reinforcement learning with environment feedback and agent-based verification grounded in real user flows.
Beyond code, the pitch is documents, spreadsheets, presentations, dashboards and meeting artifacts read directly rather than transcribed into a prompt first. The vision benchmarks back it unevenly: 62.4 on OfficeQA Pro against Opus 4.8's 48.9, but 53.4 on BabyVision where Gemini 3.7 Flash reaches 70.9. Strong on documents and charts, weaker on general visual reasoning.
It spent six days in stealth as Ox Alpha
From 20 August, OpenRouter carried a free, unnamed model with a million-token context window. It went to the top of the charts and the guessing ran for most of a week. Z.ai confirmed the identity on launch day: "Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week."
The activity graph finished on 3.94 trillion prompt tokens. We covered the reveal and who guessed right separately, and the short version is that the thread beat the press to it.

There is a cost to this method that lands on users rather than the lab. Anyone who routed proprietary code through Ox Alpha during the preview sent it to Z.ai without knowing that is where it was going. The retention policy was disclosed at the time; the recipient was not. Free previews of unnamed models are cheap in the way that matters least.
Served entirely on Chinese AI chips
This is the part of the launch with implications past the model, and Z.ai devotes a full section to it.
The entire stealth run, all 3.94 trillion tokens of it, was served on a large-scale cluster of Chinese AI accelerators rather than NVIDIA hardware. Z.ai built a dedicated inference engine on top of SGLang, and the optimisation list is specific: intra-node tensor parallelism for linear attention and the LM head, ReplaySSM, W8A8 quantization, hybrid INT8/FP8/BF16 cache quantization, and layer split. At cluster scale they run an Encode-Prefill-Decode disaggregated architecture, so multimodal encoding, prefill and decode scale as independent worker pools.
The claimed result is a 3x end-to-end serving improvement over their own baseline on the same hardware, reaching what they describe as hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs.
There is a detail in there that is easy to skim past. Z.ai says the work was accelerated by a GLM-5.3-powered infrastructure agent that helped engineers develop kernels, diagnose bottlenecks and improve the serving stack. The flagship model helped build the system that serves its cheaper sibling. Nobody outside Z.ai can verify the parity claim, and it is the single most consequential sentence in the post if it holds.
Running GLM-5.3-Flash yourself, and what it costs
The weights are on Hugging Face under MIT, in both a default and a BF16 variant, with SGLang, vLLM and TokenSpeed supported at launch and other frameworks described as coming. The safetensors index on that page counts 321B parameters, where the launch post rounds to 320B.

Through the API, Z.ai's pricing page lists $0.15 per million input tokens, $0.03 cached input, and $0.50 output. A 50% discount is running until 24:00 on 9 September 2026, Singapore time, which is where the $0.075 and $0.25 figures quoted around the web come from. Budget on the list prices and treat the discount as a window.

GLM Coding Plan subscribers get it at three times the usable quota of GLM-5.3, which for most agent work makes it the obvious default on that plan. The plan's points system and peak-hour multipliers still apply.
Where GLM-5.3-Flash actually fits
Use it as a default and reach for the flagship when a task earns it. That inverts the usual advice about flagship-and-fallback tiers, and the pricing gap is what justifies it: at a ninth of the input cost with images included and weights you can host, the burden of proof moves onto the expensive model.
The wider signal is about who is setting the cost floor. A 320B model with 18B active, trained on a new corpus, served on domestic silicon at claimed NVIDIA-comparable economics, released under MIT while the flagship it is named after stays closed. Any one of those is a normal week. Together they describe a lab optimising for reach rather than for a leaderboard position.
What none of it addresses is where a long agentic job actually lives. Browser Use and Computer Use only pay off if something is running to be driven, and a million tokens of context implies work measured in hours rather than seconds. MoClaw is a hosted cloud AI computer built for that: it stays up when your laptop doesn't, next to the tools you already use. The free trial runs three days on 1,000 credits, and a $20 subscription carries 1,000 credits a month.
FAQ
What is GLM-5.3-Flash?
Z.ai's first natively multimodal model, released 26 August 2026. It has 320 billion total parameters with 18 billion active, accepts text and images, and its weights are published on Hugging Face under an MIT licence.
Is GLM-5.3-Flash the same as GLM-5.3?
No. GLM-5.3 launched on 14 August, is text-only, and shares its base model with GLM-5.2 with all gains from post-training. GLM-5.3-Flash uses a newly trained base and a different architecture, and lists at roughly a ninth of the price. They share a version number, not a model.
How much does GLM-5.3-Flash cost?
List pricing is $0.15 per million input tokens, $0.03 cached input, and $0.50 output. A 50% discount runs until 24:00 on 9 September 2026 Singapore time, which is why $0.075 and $0.25 appear on OpenRouter and elsewhere.
Was Ox Alpha GLM-5.3-Flash?
Yes. Z.ai states in its launch post that it tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter before release, and OpenRouter's listing now carries a matching de-cloaking banner. The preview ran from 20 August and drew 3.94 trillion prompt tokens.
Can I run GLM-5.3-Flash locally?
The weights are downloadable under MIT and SGLang, vLLM and TokenSpeed support it at launch. At 320B total parameters it needs serious hardware, though only 18B parameters activate per token, which makes serving cheaper than the total size suggests.
Is GLM-5.3-Flash better than Claude Opus 4.8?
Not on Z.ai's own coding evaluation, where it reaches 29.0 against Opus 4.8's 29.5 at max effort. Z.ai's framing is "approaching" rather than beating. On several agentic benchmarks it scores higher, and it costs a fraction as much.
Specifications, benchmark figures, pricing and serving details were read from Z.ai's GLM-5.3-Flash launch post, its pricing page, the zai-org model card on Hugging Face and OpenRouter's listing on 27 August 2026. The 50% discount is scheduled to end 9 September 2026.
Continue Reading
More GuideThe MoClaw editorial team writes about workflow automation, AI agents, and the tools we build. Default byline for industry overviews, listicles, and collaborative pieces.
Ready to put this into practice?
MoClaw runs browser tasks, research, and schedules automatically. Try it free.
References: GLM-5.3-Flash: Frontier Intelligence, Flash Cost (Z.ai) · zai-org/GLM-5.3-Flash on Hugging Face · Z.ai model pricing · GLM-5.3-Flash model overview, Z.ai developer documentation · Ox Alpha listing on OpenRouter · GLM-5 technical report