Mixture of Experts

Apodex 1.1 mini - apodex 1 1 mini frontieragent framework release a relay baton lying flat on two rests

Apodex Releases the FrontierAgent Framework and the 35B Apodex 1.1 mini

Apodex released a 35B open-weight sibling to its 397B flagship, plus FrontierAgent, an Apache 2.0 agent runtime with ReAct and Agent Team modes. We work from the Hugging Face config, the GitHub repository and the arXiv report rather than the announcement — including the real parameter count, which benchmark scores belong to which model, and why the quantised builds are outrunning the originals.

Read more
deepseek v4 1 flash 552b moe model hugging face a deepseek v4.1 flash stylised whale v2

DeepSeek V4.1 Flash 552B MoE Model Released on Hugging Face. It Beats V4 Pro on Agents and Trails It on Knowledge

DeepSeek released V4.1 Flash on Hugging Face on 10 September 2026: a 552B mixture-of-experts model under the MIT licence that activates 8B parameters to read and 16B to write, stores its global KV cache in 890 bytes per token, and replaces V4 Pro for every API caller from 04:00 UTC on 14 September. Its own model card shows it leading V4 Pro on all 12 agentic rows and trailing it on 12 of 16 base-model rows, including a 12.9-point gap on SimpleQA-Verified. This breakdown covers the causal encoder-decoder architecture, the KV cache arithmetic, the new price sheet with a worked agent bill, the 511 GB checkpoint and 614 GB serving floor, the mismatched reasoning-effort aliases, and a pre-cutover checklist.

Read more
2-bit quantization - 2 bit quantization glm 5 3 flash macbook pro a shallow open tray with one block resting inside it

2-Bit Quantization Puts GLM-5.3-Flash on a Laptop. The Arithmetic Says Which One

2-bit quantization is the only reason a 320-billion-parameter model appears in the same sentence as a laptop. Z.ai shipped GLM-5.3-Flash in August 2026 under an MIT licence, and within days the community had squeezed the full-precision checkpoint from 641.64 GB down to 108.72 GB. The headlines that followed said the same thing in different words: […]

Read more
GLM-5.3-Flash - glm 5 3 flash open weight 320b model a solid lightning bolt

GLM-5.3-Flash: The 320B Open-Weight Model That Ran on Chinese Chips

Z.ai spent six days serving an anonymous model called Ox Alpha on OpenRouter, took nearly 20% of the platform’s weekly token share, and only then revealed it was GLM-5.3-Flash — a 320B mixture-of-experts model with 18B active parameters, a one-million-token context window and MIT-licensed weights. This piece works through the hybrid attention architecture, what the benchmark table supports and what it does not, what the API actually costs once the launch promotion ends, how credible the domestic-silicon claim is, and what any of it changes for a business choosing a model this quarter.

Read more
Qwen3.8-Flash - qwen3 8 flash next 125b moe model a honeycomb block seven cells

Alibaba Releases Qwen3.8-Flash: A Multimodal 125B MoE Model That Previews Qwen4

Alibaba open-weighted Qwen3.8-Flash-Next on 26 August 2026: a multimodal mixture-of-experts model with 125 billion parameters, a separate 51-billion-parameter N-gram embedding table, and just 6 billion parameters activated per token. This breakdown covers the four rebuilt subsystems — Gated DeltaNet paired with Qwen Sparse Attention at block granularity, a Gated Residual stream widened to four gated branches, the N-gram table that offloads to host RAM, and the Muon plus AdamW training recipe with batch-size warmup removed — alongside the 48-layer stack of 512 experts that fires eleven per token, the published benchmark table showing 62.5 on SWE-bench Pro against 53.4 for Claude Opus 4.6 and 84.5 on AndroidWorld against 62.0, the single loss on Humanity’s Last Exam at 35.9 against 40.0, the unverifiable one-ninth training cost claim, the 262,144-token native context extended to a million with YaRN, hosted pricing of $0.16 and $0.47 per million tokens against $2.00 and $6.00 for Qwen3.8-Max, the real hardware bill from a 172.78 GiB FP8 checkpoint down to a 111 GB four-bit GGUF, the qwen-community-1.0 licence that is not Apache 2.0, and a buyer’s checklist for treating a preview checkpoint as a production dependency.

Read more
Z.ai - z ai lab behind ox alpha model a treasure chest closed lid

Surprise: Z.ai Is the AI Lab Behind the Mysterious Ox Alpha Model

On 26 August 2026 the mystery ended: Z.ai, the Beijing lab formerly known as Zhipu AI, confirmed that the anonymous Ox Alpha model topping OpenRouter and OpenCode was the newest iteration of its GLM series, and published the weights the same evening as GLM-5.3-Flash under an MIT licence. This breakdown covers what was confirmed and when, the architecture the model card revealed — 320 billion total parameters with just 18 billion active, hybrid sparse and linear attention, a 1,048,576-token context and forced reasoning that cannot be disabled — the published benchmark table showing 84.3 on Terminal Bench 2.1 against 85.0 for Claude Opus 4.8 and 87.4 for GPT-5.6 Terra, the viral 80 per cent DeepSWE claim that came from a 10-task subset and collapsed to 63.4 on the full 113-task run, the 44 trillion tokens and 503,000 users the stealth week generated, the tokenizer and error-code forensics that unmasked the lab before it spoke, the $0.15 and $0.50 per million token pricing, the company’s Hong Kong listing and US entity-list status, and a buyer’s checklist for deciding between the hosted API and self-hosted weights.

Read more
deepseek v4 complete guide a three ascending rounded pillars

DeepSeek V4 Complete Guide: Best Open-Weight AI of 2026

DeepSeek V4 is the MIT-licensed open-weight family that replaced the never-released R2, pairing a one-million-token context window with sub-dollar output pricing. This complete guide covers the V4 Pro 0813 and V4 Flash 0731 GA builds, their official benchmarks, the peak/off-peak billing change landing on 16 August 2026, and a practical framework for choosing between the two models.

Read more
CHAT