Beijing’s Zhipu AI has released its premium GLM-5.3-Flash model claiming it runs entirely on roughly 100,000 domestically manufactured Chinese chips. In the process it confirmed that the anonymous “Ox” model which topped OpenRouter’s usage charts earlier this month was its own.
Z.AI (Zhipu AI) unveiled GLM-5.3-Flash on August 26 and 27. It is a 320-billion-parameter model with a one-million-token context window, released under an MIT open-source license. It is priced at around US$0.15 per million input tokens and US$0.50 per million output tokens on OpenRouter.
That pricing undercuts Western frontier models by a large margin. Anthropic’s Opus 5 charges US$5 per million input tokens and US$25 per million output tokens, so GLM-5.3-Flash runs roughly 33 times cheaper on input and 50 times cheaper on output.
The headline claim is architectural as well as commercial. The company says the model is fully powered by roughly 100,000 China-made AI accelerators and is architected for deployment across domestic Chinese GPU families. It is a pointed rebuttal to the narrative that state-of-the-art Chinese foundational AI still depends on Nvidia hardware.
Alongside the launch, Z.AI confirmed that the mysterious “Ox Alpha” model was in fact its own GLM-5.3-Flash running on Chinese GPUs. Ox Alpha was an anonymised “stealth” offering that briefly topped OpenRouter’s model-usage rankings in mid-August before being attributed to a third-party provider. The reveal gives the launch an unusually sharp narrative arc. An event that English-language commentators had flagged as an unknown new entrant was, a week later, an open-source release from a recognised Beijing lab.
Markets noticed. Z.ai shares gained about 8% on the announcement. Caixin, which first reported on the anonymous OpenRouter usage behind the model, confirmed the development alongside CNBC, Global Times and US tech site Wccftech.
The release underscores a wider trend. Beijing AI labs are making visible progress in running advanced models on domestic silicon even as US export restrictions limit access to high-end Nvidia chips. GLM-5.3-Flash is engineered to run across multiple domestic accelerator lines, not just Nvidia’s CUDA.
