agent benchmarks

evoharness meta 8b model match claude opus 4 5 a two cubes small and large

Meta Researchers Taught an 8B AI Model to Match Claude Opus 4.5

Meta AI and the University of Illinois have published EvoHarness-RL, a framework that trains an agent to run its own memory, progress and experience stores instead of following a hand-written harness. On the ALFWorld benchmark it took Alibaba’s open-weight Qwen3-8B from 47.9% to 96.9%, a whisker past Claude Opus 4.5’s unaided 96.4%. This breakdown covers what the framework actually trains, the full results table rather than the two rows that travel well, why the same harness pushes Opus 4.5 to 98.5%, why Claude Opus sits inside the training loop as teacher and consolidator, what harness annealing and harness evolution mean for latency, and what any of it is worth if you are shipping agents commercially.

Read more
CHAT