LLM pricing

deepseek v4 1 flash 552b moe model hugging face a deepseek v4.1 flash stylised whale v2

DeepSeek V4.1 Flash 552B MoE Model Released on Hugging Face. It Beats V4 Pro on Agents and Trails It on Knowledge

DeepSeek released V4.1 Flash on Hugging Face on 10 September 2026: a 552B mixture-of-experts model under the MIT licence that activates 8B parameters to read and 16B to write, stores its global KV cache in 890 bytes per token, and replaces V4 Pro for every API caller from 04:00 UTC on 14 September. Its own model card shows it leading V4 Pro on all 12 agentic rows and trailing it on 12 of 16 base-model rows, including a 12.9-point gap on SimpleQA-Verified. This breakdown covers the causal encoder-decoder architecture, the KV cache arithmetic, the new price sheet with a worked agent bill, the 511 GB checkpoint and 614 GB serving floor, the mismatched reasoning-effort aliases, and a pre-cutover checklist.

Read more
gemini 3 8 flash aixploria listing a ballot box lid slot

Gemini 3.8 Flash Is Now Available on AIxploria

AIxploria listing pages are how a very large audience meets a new model for the first time, and Gemini 3.8 Flash got one at 03:16 UTC on 3 September 2026, roughly a day after Google shipped the model itself. The card sits at 4.4 out of 5 stars, carries 84 upvotes, and files Google’s newest […]

Read more
claude fable 5 1 aixploria a claude fable 5.1 lighthouse tower

New AI Tool: Claude Fable 5.1 Has Just Landed on AIxploria

Claude Fable 5.1 is now listed on AIxploria, the AI tools directory that catalogues thousands of AI sites across more than fifty categories. The new card sits at number one on the directory’s “Latest AI” feed, carries a Gold Verified badge, is priced “Paid”, and files Anthropic’s frontier model under LLM models and SuperTools. For […]

Read more
GLM-5.3-Flash - glm 5 3 flash open weight 320b model a solid lightning bolt

GLM-5.3-Flash: The 320B Open-Weight Model That Ran on Chinese Chips

Z.ai spent six days serving an anonymous model called Ox Alpha on OpenRouter, took nearly 20% of the platform’s weekly token share, and only then revealed it was GLM-5.3-Flash — a 320B mixture-of-experts model with 18B active parameters, a one-million-token context window and MIT-licensed weights. This piece works through the hybrid attention architecture, what the benchmark table supports and what it does not, what the API actually costs once the launch promotion ends, how credible the domestic-silicon claim is, and what any of it changes for a business choosing a model this quarter.

Read more
Qwen3.8-Flash - qwen3 8 flash next 125b moe model a honeycomb block seven cells

Alibaba Releases Qwen3.8-Flash: A Multimodal 125B MoE Model That Previews Qwen4

Alibaba open-weighted Qwen3.8-Flash-Next on 26 August 2026: a multimodal mixture-of-experts model with 125 billion parameters, a separate 51-billion-parameter N-gram embedding table, and just 6 billion parameters activated per token. This breakdown covers the four rebuilt subsystems — Gated DeltaNet paired with Qwen Sparse Attention at block granularity, a Gated Residual stream widened to four gated branches, the N-gram table that offloads to host RAM, and the Muon plus AdamW training recipe with batch-size warmup removed — alongside the 48-layer stack of 512 experts that fires eleven per token, the published benchmark table showing 62.5 on SWE-bench Pro against 53.4 for Claude Opus 4.6 and 84.5 on AndroidWorld against 62.0, the single loss on Humanity’s Last Exam at 35.9 against 40.0, the unverifiable one-ninth training cost claim, the 262,144-token native context extended to a million with YaRN, hosted pricing of $0.16 and $0.47 per million tokens against $2.00 and $6.00 for Qwen3.8-Max, the real hardware bill from a 172.78 GiB FP8 checkpoint down to a 111 GB four-bit GGUF, the qwen-community-1.0 licence that is not Apache 2.0, and a buyer’s checklist for treating a preview checkpoint as a production dependency.

Read more
deepseek v4 complete guide a three ascending rounded pillars

DeepSeek V4 Complete Guide: Best Open-Weight AI of 2026

DeepSeek V4 is the MIT-licensed open-weight family that replaced the never-released R2, pairing a one-million-token context window with sub-dollar output pricing. This complete guide covers the V4 Pro 0813 and V4 Flash 0731 GA builds, their official benchmarks, the peak/off-peak billing change landing on 16 August 2026, and a practical framework for choosing between the two models.

Read more
GPT-5.6 Sol - gpt 5 6 sol a three ascending rounded pillars

GPT-5.6 Sol: Complete Guide to OpenAI’s Best Model Yet

GPT-5.6 Sol is OpenAI’s flagship model, launched publicly on 9 July 2026 alongside Terra and Luna. This guide maps the tier scheme, the full API price list including the 272K long-context surcharge, honest benchmarks against Anthropic’s Claude, and availability across ChatGPT plans and cloud platforms. It closes with what the Doug pre-training project and Astra signal about GPT-6.

Read more
CHAT