multimodal AI

deepseek v4 1 flash 552b moe model hugging face a deepseek v4.1 flash stylised whale v2

DeepSeek V4.1 Flash 552B MoE Model Released on Hugging Face. It Beats V4 Pro on Agents and Trails It on Knowledge

DeepSeek released V4.1 Flash on Hugging Face on 10 September 2026: a 552B mixture-of-experts model under the MIT licence that activates 8B parameters to read and 16B to write, stores its global KV cache in 890 bytes per token, and replaces V4 Pro for every API caller from 04:00 UTC on 14 September. Its own model card shows it leading V4 Pro on all 12 agentic rows and trailing it on 12 of 16 base-model rows, including a 12.9-point gap on SimpleQA-Verified. This breakdown covers the causal encoder-decoder architecture, the KV cache arithmetic, the new price sheet with a worked agent bill, the 511 GB checkpoint and 614 GB serving floor, the mismatched reasoning-effort aliases, and a pre-cutover checklist.

Read more
GLM-5.3-Flash - glm 5 3 flash open weight 320b model a solid lightning bolt

GLM-5.3-Flash: The 320B Open-Weight Model That Ran on Chinese Chips

Z.ai spent six days serving an anonymous model called Ox Alpha on OpenRouter, took nearly 20% of the platform’s weekly token share, and only then revealed it was GLM-5.3-Flash — a 320B mixture-of-experts model with 18B active parameters, a one-million-token context window and MIT-licensed weights. This piece works through the hybrid attention architecture, what the benchmark table supports and what it does not, what the API actually costs once the launch promotion ends, how credible the domestic-silicon claim is, and what any of it changes for a business choosing a model this quarter.

Read more
Omni 1.1 - google omni 1 1 flash video model 1080p 4k a upright panel raised play triangle

Google Releases Omni 1.1 Flash Video Model With 1080p and 4K Support

Google released Gemini Omni 1.1 Flash on 27 August 2026, the production version of its generative video model, and the headline change is output above 720p. The model now accepts resolution values of 360p, 720p, 1080p and 4K, extends an existing clip in ten-second increments to a cumulative forty seconds, and interpolates footage between a start frame and an end frame you supply. It is paid-tier only, billed on output tokens at a rate Google states as 5,792 tokens per second of 720p video — roughly $0.10 per second — with a 360p draft tier at a third of that cost and up to 60% faster. This breakdown covers exactly what shipped, why the 1080p and 4K figures are documented as upscales rather than native renders, how the token billing compares against Veo 3.1, which surfaces the model is live on today, and the limitations list that blocks editing uploaded video for users in the UK, the EEA and Switzerland.

Read more
Qwen3.8-Flash - qwen3 8 flash next 125b moe model a honeycomb block seven cells

Alibaba Releases Qwen3.8-Flash: A Multimodal 125B MoE Model That Previews Qwen4

Alibaba open-weighted Qwen3.8-Flash-Next on 26 August 2026: a multimodal mixture-of-experts model with 125 billion parameters, a separate 51-billion-parameter N-gram embedding table, and just 6 billion parameters activated per token. This breakdown covers the four rebuilt subsystems — Gated DeltaNet paired with Qwen Sparse Attention at block granularity, a Gated Residual stream widened to four gated branches, the N-gram table that offloads to host RAM, and the Muon plus AdamW training recipe with batch-size warmup removed — alongside the 48-layer stack of 512 experts that fires eleven per token, the published benchmark table showing 62.5 on SWE-bench Pro against 53.4 for Claude Opus 4.6 and 84.5 on AndroidWorld against 62.0, the single loss on Humanity’s Last Exam at 35.9 against 40.0, the unverifiable one-ninth training cost claim, the 262,144-token native context extended to a million with YaRN, hosted pricing of $0.16 and $0.47 per million tokens against $2.00 and $6.00 for Qwen3.8-Max, the real hardware bill from a 172.78 GiB FP8 checkpoint down to a 111 GB four-bit GGUF, the qwen-community-1.0 licence that is not Apache 2.0, and a buyer’s checklist for treating a preview checkpoint as a production dependency.

Read more
A beam of light passing through a translucent veil and resolving into one solid shape as ghost duplicates fade, AI sensory hallucinations dissolving into a true perception

AI Reduces Sensory Hallucinations, Even at Night or in Smoke: Inside KAIST’s DNA and MAD Methods

KAIST has published two methods for cutting sensory hallucinations in multimodal AI: the failure where a model misreads what a sensor physically reports, or invents a perception in one channel because another channel suggested it. DNA optimisation teaches vision-language models the physics of thermal, depth and X-ray sensors using their own wrong answers as the training signal. MAD suppresses cross-modal interference at decoding time with no retraining at all. Here is what each method fixes, what the reported numbers do and do not establish, where sensory hallucinations cost the most in production, and what this line of work still leaves unsolved.

Read more
CHAT