small language models

evoharness meta 8b model match claude opus 4 5 a two cubes small and large

Meta Researchers Taught an 8B AI Model to Match Claude Opus 4.5

Meta AI and the University of Illinois have published EvoHarness-RL, a framework that trains an agent to run its own memory, progress and experience stores instead of following a hand-written harness. On the ALFWorld benchmark it took Alibaba’s open-weight Qwen3-8B from 47.9% to 96.9%, a whisker past Claude Opus 4.5’s unaided 96.4%. This breakdown covers what the framework actually trains, the full results table rather than the two rows that travel well, why the same harness pushes Opus 4.5 to 98.5%, why Claude Opus sits inside the training loop as teacher and consolidator, what harness annealing and harness evolution mean for latency, and what any of it is worth if you are shipping agents commercially.

Read more
small language models business a three ascending rounded pillars

Small Language Models: Complete 2026 Guide for Smart Teams

Small language models now handle most routine business AI work at a fraction of frontier-model cost. This guide maps the 2026 field — Phi-4-reasoning-vision, Gemma 4, Qwen 3.5 Small, Claude Haiku 4.5 and Ministral 3 — and shows exactly when a compact model beats a frontier LLM on cost, privacy and latency. It closes with a hardware plan and a 90-day deployment roadmap.

Read more
CHAT