Z.ai

GLM-5.3-Flash - glm 5 3 flash open weight 320b model a solid lightning bolt

GLM-5.3-Flash: The 320B Open-Weight Model That Ran on Chinese Chips

Z.ai spent six days serving an anonymous model called Ox Alpha on OpenRouter, took nearly 20% of the platform’s weekly token share, and only then revealed it was GLM-5.3-Flash — a 320B mixture-of-experts model with 18B active parameters, a one-million-token context window and MIT-licensed weights. This piece works through the hybrid attention architecture, what the benchmark table supports and what it does not, what the API actually costs once the launch promotion ends, how credible the domestic-silicon claim is, and what any of it changes for a business choosing a model this quarter.

Read more
Z.ai - z ai lab behind ox alpha model a treasure chest closed lid

Surprise: Z.ai Is the AI Lab Behind the Mysterious Ox Alpha Model

On 26 August 2026 the mystery ended: Z.ai, the Beijing lab formerly known as Zhipu AI, confirmed that the anonymous Ox Alpha model topping OpenRouter and OpenCode was the newest iteration of its GLM series, and published the weights the same evening as GLM-5.3-Flash under an MIT licence. This breakdown covers what was confirmed and when, the architecture the model card revealed — 320 billion total parameters with just 18 billion active, hybrid sparse and linear attention, a 1,048,576-token context and forced reasoning that cannot be disabled — the published benchmark table showing 84.3 on Terminal Bench 2.1 against 85.0 for Claude Opus 4.8 and 87.4 for GPT-5.6 Terra, the viral 80 per cent DeepSWE claim that came from a 10-task subset and collapsed to 63.4 on the full 113-task run, the 44 trillion tokens and 503,000 users the stealth week generated, the tokenizer and error-code forensics that unmasked the lab before it spoke, the $0.15 and $0.50 per million token pricing, the company’s Hong Kong listing and US entity-list status, and a buyer’s checklist for deciding between the hosted API and self-hosted weights.

Read more
CHAT