LLM-as-a-judge

ai benchmarks tests grade ai wrong a carnival high striker with a bell and mallet

The Tests That Grade AI May Be Getting It Wrong

Two Stanford-led studies for COLM 2026 test AI benchmarks the way psychologists test exams. Across 56 benchmarks and 53 models, safety tests with the same label often disagreed, capability tests blurred together, and the scoring format mattered more than the topic. We explain the findings and what they mean for choosing a model.

Read more
CHAT