The Tests That Grade AI May Be Getting It Wrong
Two Stanford-led studies for COLM 2026 test AI benchmarks the way psychologists test exams. Across 56 benchmarks and 53 models, safety tests with the same label often disagreed, capability tests blurred together, and the scoring format mattered more than the topic. We explain the findings and what they mean for choosing a model.