N-gram Embedding

Qwen3.8-Flash - qwen3 8 flash next 125b moe model a honeycomb block seven cells

Alibaba Releases Qwen3.8-Flash: A Multimodal 125B MoE Model That Previews Qwen4

Alibaba open-weighted Qwen3.8-Flash-Next on 26 August 2026: a multimodal mixture-of-experts model with 125 billion parameters, a separate 51-billion-parameter N-gram embedding table, and just 6 billion parameters activated per token. This breakdown covers the four rebuilt subsystems — Gated DeltaNet paired with Qwen Sparse Attention at block granularity, a Gated Residual stream widened to four gated branches, the N-gram table that offloads to host RAM, and the Muon plus AdamW training recipe with batch-size warmup removed — alongside the 48-layer stack of 512 experts that fires eleven per token, the published benchmark table showing 62.5 on SWE-bench Pro against 53.4 for Claude Opus 4.6 and 84.5 on AndroidWorld against 62.0, the single loss on Humanity’s Last Exam at 35.9 against 40.0, the unverifiable one-ninth training cost claim, the 262,144-token native context extended to a million with YaRN, hosted pricing of $0.16 and $0.47 per million tokens against $2.00 and $6.00 for Qwen3.8-Max, the real hardware bill from a 172.78 GiB FP8 checkpoint down to a 111 GB four-bit GGUF, the qwen-community-1.0 licence that is not Apache 2.0, and a buyer’s checklist for treating a preview checkpoint as a production dependency.

Read more
CHAT