For six days in August, the most-used AI model on OpenRouter had no name attached to it. Developers called it Ox Alpha, and nobody, not even OpenRouter itself, would say who built it. On August 26, the mystery ended: Z.ai confirmed Ox Alpha was actually an early preview of GLM-5.3-Flash, and released the full weights on Hugging Face the same night. Here’s what actually happened, and how to use the model now that it has a real name.
Ox Alpha was a stealth preview of GLM-5.3-Flash, a free multimodal coding model from Chinese lab Z.ai. The free preview has ended, but the real model is now live everywhere as z-ai/glm-5.3-flash, priced from $0.075 per million input tokens through September 9 under a launch promotion.
How the Mystery Unfolded
OpenRouter listed Ox Alpha on August 20 as a “stealth model,” free to use, with a roughly 1-million-token context window and no attribution beyond “an unnamed third-party provider.” Within a day, developer Ben Davis ran it against 10 tasks from the DeepSWE benchmark and reported 80% first-pass accuracy, ahead of Claude Fable 5 at 65% and GPT-5.6 Sol at 52%. That single data point sent the model viral and knocked DeepSeek off the top of OpenCode’s leaderboard.
Speculation immediately centered on Chinese labs. Developers fingerprinting the model’s tokenizer found it matched Z.ai’s GLM family on nearly every technical probe, while other analysts floated Xiaomi’s MiMo team as an alternative. Bloomberg reported that coding tools, including Claude Code, routed billions of tokens through the anonymous model in less than a week. Then, on August 26, Z.ai settled the debate itself by publishing GLM-5.3-Flash openly rather than letting the leak continue.
What GLM-5.3-Flash Actually Is
| Spec | Detail |
|---|---|
| Parameters | 320B total, 18B active (Mixture-of-Experts) |
| Context window | Just over 1 million tokens |
| Inputs supported | Text, image, and video |
| Architecture | Hybrid sparse + linear attention |
| License | MIT (open weights, commercial use allowed) |
Zhipu AI, the company now branded Z.ai, was added to the US Commerce Department’s Entity List in January 2025 over its role in China’s military technology development. This doesn’t block API access for most individual developers, but it’s a relevant fact for companies doing vendor due diligence.
How to Use GLM-5.3-Flash Now
Step 1: Try It in the Browser
Go to chat.z.ai and start a conversation directly — GLM-5.3-Flash now powers Z.ai’s own chatbot by default, no setup required.
Step 2: Get API Access
Create an account through Z.ai’s developer platform, then reference the model as glm-5.3-flash in your API calls. Through September 9, pricing sits at a 50% launch discount: $0.075 per million input tokens and $0.25 per million output tokens, reverting to $0.15 / $0.50 afterward.
Step 3: Route Through OpenRouter Instead (Optional)
If you’re already set up on OpenRouter, the model is listed there as z-ai/glm-5.3-flash, giving you access through the same API key you use for other models, with automatic routing across multiple hosting providers.
Step 4: Self-Host the Open Weights (Advanced)
Because the weights are MIT-licensed and public on Hugging Face, you can technically self-host GLM-5.3-Flash. However, at 320 billion total parameters, realistic self-hosting requires serious GPU infrastructure, not a personal workstation.
Any code still pointed at the old stealth identifier stealth/ox-alpha needs updating — that listing is gone from OpenRouter’s catalog now that the free preview has ended.
Frequently Asked Questions
Is GLM-5.3-Flash still free?
No. The free anonymous preview under the Ox Alpha name has ended. The model is now paid, though the current launch discount makes it one of the cheapest capable coding models available.
Was the viral DeepSWE benchmark result official?
No — it came from one developer’s informal 10-task test, not a rigorous, peer-reviewed benchmark. It’s a useful early signal, but worth verifying against your own workload rather than treating it as a guaranteed result.
Why did Z.ai release the model anonymously in the first place?
Running a free, unbranded preview let Z.ai collect real usage data and unbiased benchmark chatter before the model’s origin colored how people judged it. Several other labs, including ByteDance and Alibaba, have used the same anonymous-launch tactic this year.