DeepSeek V4.1 Flash

deepseek v4 1 flash 552b moe model hugging face a deepseek v4.1 flash stylised whale v2

DeepSeek V4.1 Flash 552B MoE Model Released on Hugging Face. It Beats V4 Pro on Agents and Trails It on Knowledge

DeepSeek released V4.1 Flash on Hugging Face on 10 September 2026: a 552B mixture-of-experts model under the MIT licence that activates 8B parameters to read and 16B to write, stores its global KV cache in 890 bytes per token, and replaces V4 Pro for every API caller from 04:00 UTC on 14 September. Its own model card shows it leading V4 Pro on all 12 agentic rows and trailing it on 12 of 16 base-model rows, including a 12.9-point gap on SimpleQA-Verified. This breakdown covers the causal encoder-decoder architecture, the KV cache arithmetic, the new price sheet with a worked agent bill, the 511 GB checkpoint and 614 GB serving floor, the mismatched reasoning-effort aliases, and a pre-cutover checklist.

Read more
CHAT