AI benchmarks

GLM-5.3-Flash - glm 5 3 flash open weight 320b model a solid lightning bolt

GLM-5.3-Flash: The 320B Open-Weight Model That Ran on Chinese Chips

Z.ai spent six days serving an anonymous model called Ox Alpha on OpenRouter, took nearly 20% of the platform’s weekly token share, and only then revealed it was GLM-5.3-Flash — a 320B mixture-of-experts model with 18B active parameters, a one-million-token context window and MIT-licensed weights. This piece works through the hybrid attention architecture, what the benchmark table supports and what it does not, what the API actually costs once the launch promotion ends, how credible the domestic-silicon claim is, and what any of it changes for a business choosing a model this quarter.

Read more
Z.ai - z ai lab behind ox alpha model a treasure chest closed lid

Surprise: Z.ai Is the AI Lab Behind the Mysterious Ox Alpha Model

On 26 August 2026 the mystery ended: Z.ai, the Beijing lab formerly known as Zhipu AI, confirmed that the anonymous Ox Alpha model topping OpenRouter and OpenCode was the newest iteration of its GLM series, and published the weights the same evening as GLM-5.3-Flash under an MIT licence. This breakdown covers what was confirmed and when, the architecture the model card revealed — 320 billion total parameters with just 18 billion active, hybrid sparse and linear attention, a 1,048,576-token context and forced reasoning that cannot be disabled — the published benchmark table showing 84.3 on Terminal Bench 2.1 against 85.0 for Claude Opus 4.8 and 87.4 for GPT-5.6 Terra, the viral 80 per cent DeepSWE claim that came from a 10-task subset and collapsed to 63.4 on the full 113-task run, the 44 trillion tokens and 503,000 users the stealth week generated, the tokenizer and error-code forensics that unmasked the lab before it spoke, the $0.15 and $0.50 per million token pricing, the company’s Hong Kong listing and US entity-list status, and a buyer’s checklist for deciding between the hosted API and self-hosted weights.

Read more
ox alpha mystery ai model a solid theatre mask

0x Alpha: The Mystery AI Model Everyone Is Testing for Free

Ox Alpha — often searched as 0x Alpha — appeared anonymously on OpenRouter on 20 August 2026: a free frontier-class reasoning model with a million-token context window, video input and a viral claim to beat paid rivals at coding. This guide covers the verified specification sheet, the 10-task benchmark caveats, the serving-layer forensics pointing to Zhipu AI, the four previous stealth models that were all claimed by Chinese labs, what the free week really costs in data terms on each of the two access routes, and a defensive pilot plan for businesses that want the free market intelligence without handing an anonymous operator their company data.

Read more
Gemini 3.1 Pro - gemini 3 1 pro a three ascending rounded pillars

Gemini 3.1 Pro: Complete Guide to Google’s Best AI Model

Gemini 3.1 Pro remains Google’s flagship AI model in August 2026, six months after its February preview launch. This complete guide covers its 1M-token context window, benchmark scores, API pricing and every access route. It also explains how the flagship compares with the fast-moving Gemini Flash line and the still-delayed Gemini 3.5 Pro.

Read more
deepseek v4 complete guide a three ascending rounded pillars

DeepSeek V4 Complete Guide: Best Open-Weight AI of 2026

DeepSeek V4 is the MIT-licensed open-weight family that replaced the never-released R2, pairing a one-million-token context window with sub-dollar output pricing. This complete guide covers the V4 Pro 0813 and V4 Flash 0731 GA builds, their official benchmarks, the peak/off-peak billing change landing on 16 August 2026, and a practical framework for choosing between the two models.

Read more
GPT-5.6 Sol - gpt 5 6 sol a three ascending rounded pillars

GPT-5.6 Sol: Complete Guide to OpenAI’s Best Model Yet

GPT-5.6 Sol is OpenAI’s flagship model, launched publicly on 9 July 2026 alongside Terra and Luna. This guide maps the tier scheme, the full API price list including the 272K long-context surcharge, honest benchmarks against Anthropic’s Claude, and availability across ChatGPT plans and cloud platforms. It closes with what the Doug pre-training project and Astra signal about GPT-6.

Read more
CHAT