AI benchmarks are the standard tests that rank models on reasoning, knowledge, safety and bias, and they decide which model looks best in a launch post, a leaderboard or a procurement shortlist. Two new studies led by Stanford researchers ask a blunt question: do those tests measure what their names say? Often, they conclude, they do not.
The studies will be presented in October at the Conference on Language Modeling (COLM) in San Francisco, a leading venue for research on natural language processing. Stanford’s Institute for Human-Centered AI (HAI) published an interview with two of the authors, assistant professor Sanmi Koyejo and graduate student Sang Truong, on 25 September 2026, and the story was also carried by Stanford Report and Tech Xplore under the headline “The tests that grade AI may be getting it wrong”. “Whoever holds the measuring stick steers the ship,” Koyejo said.
This article explains what the researchers did, the four main findings, why safety scores break down across languages, and what anyone choosing or buying a model should do differently. The figures come from the two papers; where we add arithmetic, we say so. For practical testing advice, see our guides to AI agent evaluation metrics and to the enterprise AI evaluation gap.
Table of contents
- What the Stanford Research Found About AI Benchmarks
- How the Researchers Tested AI Benchmarks
- Finding One: Safety AI Benchmarks Often Disagree
- Finding Two: Capability AI Benchmarks Blur Together
- Finding Three: The Format of AI Benchmarks Beats the Topic
- Finding Four: A Bias Test That Measures Reasoning
- Why Safety AI Benchmarks Break Across Languages
- The Problems Grow With Newer Models
- What This Means for Anyone Choosing a Model With AI Benchmarks
- Limits of the AI Benchmarks Research
- AI Benchmarks FAQ
- References and Further Reading
What the Stanford Research Found About AI Benchmarks
The two papers look at different problems but share a method: treat a benchmark as a measuring instrument and test it the way psychologists test an exam.
The 56-benchmark study
The larger paper is titled “What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks”. It was posted to arXiv on 8 September 2026. Its 11 authors, with Meera Desai of the University of Michigan as first author, come from Stanford, Michigan, Microsoft Research, Cornell Tech, Yale and the health AI company Abridge.
The multilingual safety study
The second paper, “Why Do Safety Guardrails Degrade Across Languages?”, is by Max Zhang, Ameen Patel, Sang T. Truong and Sanmi Koyejo. It was first posted in May 2026 and revised in August, and is a COLM 2026 poster. It asks why safety tests give worse results in languages other than English, and whether the test or the model is to blame.
| Paper | Scale | Headline finding |
|---|---|---|
| What AI Benchmarks Actually Measure | 56 benchmarks, 53 models | Many benchmarks do not measure their stated concept; score format predicts agreement better than topic |
| Why Do Safety Guardrails Degrade Across Languages? | 61 model configurations, 10 languages, 1.9 million responses | A single jailbreak rate hides several causes; 22 configurations were more vulnerable in English than in low-resource languages |
How the Researchers Tested AI Benchmarks
Koyejo compares the approach to standardised testing such as the SAT. “Borrowing tools from measurement science and psychometrics – the century-old science of measuring human abilities – our research shows that benchmarks do not always measure what they claim to,” he told HAI.
Two checks borrowed from psychology
The first check is convergent validity: tests that claim to measure the same thing should agree. If two reasoning benchmarks rank models very differently, at least one of them is measuring something else. The second is discriminant validity: tests that claim to measure different things should disagree more. If a bias test ranks models exactly like a reasoning test, it may be a reasoning test in disguise.
Grouping benchmarks by what they claim
Because few AI benchmarks have ever been validated, the team did not compare each new test to one trusted “gold standard”. Instead, they grouped benchmarks with similar stated goals under a shared label, such as reasoning, knowledge, bias or refusal, and asked whether the members of each group ranked the models in the same order. They measured agreement with Spearman rank correlation, and also fitted item-level statistical models to see whether one underlying skill could explain the answers.
Cleaning the data first
The team collected outputs from 53 models on 56 benchmarks and released the dataset publicly. They dropped four benchmarks as “saturated”, because the best model and the median model were too close to tell apart: DecodingTrust Stereotype, CIVICS, BoolQ and IMDB. They dropped four more whose free-text answers were marked by exact matching, after finding the scores depended on formatting rather than substance. That left 48 benchmarks for the main analysis.
A check against HELM
As a sanity test, the team compared its model rankings with those from Stanford’s HELM project on shared benchmarks. On 10 benchmarks the average rank correlation was 0.74, a strong match. On the four exact-match free-text benchmarks it was only 0.25, which is why those four were excluded.
How the 56 AI benchmarks were filtered to 48 (paper, Appendix A)
Finding One: Safety AI Benchmarks Often Disagree
The first result concerns tests that share a label. Capability benchmarks labelled with the same concept, such as knowledge, generally agreed on how to rank models. Safety benchmarks did not.
Refusal, safety detection and bias
The paper reports that safety AI benchmarks labelled refusal, safety detection and bias “show wide interquartile ranges and correlations that frequently approach or fall below zero”. A correlation near zero means knowing a model’s rank on one test tells you almost nothing about its rank on another test with the same label.
Not necessarily bad tests
The authors are careful here. “We do not interpret lower correlations among safety benchmarks as evidence of poor construction,” they write, “as many safety concepts may be inherently multi-dimensional.” Refusing a request for weapons instructions and refusing a request for medical advice are both “refusal”, but they may be different skills. The practical point stands: one safety number cannot stand in for safety in general.
Finding Two: Capability AI Benchmarks Blur Together
The second result is the mirror image. Tests that claim to measure different capabilities often agreed as strongly as tests that claim to measure the same one.
Reasoning and knowledge look like one skill
Correlations between benchmarks labelled reasoning and benchmarks labelled knowledge “are often as high as correlations between model scores on benchmarks with the same assigned concept”, the paper says, with summarisation the one exception. At the item level, a single shared skill predicted answers on reasoning and knowledge tests about as well as two separate ones did. In plain terms, today’s AI benchmarks may be ranking general capability under several different names.
Some safety tests track capability
Some safety labels behaved the same way. Benchmarks labelled ethics, bias, privacy and unsafe behaviour correlated more strongly with capability benchmarks than with each other, “suggesting that some of these benchmarks may actually be measuring capability concepts”. The exception was a pair designed to pull apart: refusal and over-refusal tests were strongly inversely related, as intended, since a model that refuses more will score worse on tests of refusing too much.
Finding Three: The Format of AI Benchmarks Beats the Topic
The most striking result is about how a test is marked, not what it asks.
Multiple choice, exact match and AI judges
The team grouped benchmarks by score format: multiple-choice questions, free-text answers marked by exact matching, and free-text answers marked by another AI model acting as a judge. A refusal benchmark written as multiple choice, SGBench-mcq, ranked models like other multiple-choice tests rather than like other refusal tests. Benchmarks marked by an AI judge formed their own distinct cluster.
The numbers behind it
A statistical test (a partial Mantel test) weighed shared format against shared concept as predictors of how closely two benchmarks agree. With format coded in three groups, format had about twice the effect of concept: 0.275 against 0.138. When format was coded simply as “AI judge or not”, format rose to 0.526 and concept fell to −0.058, which was not statistically significant. “Shared use of LLM-judge scoring is a stronger predictor of benchmark similarity than shared concept,” the authors write.
Effect on agreement between two benchmarks: shared format versus shared concept (partial Mantel coefficients, paper section 4.3)
Same benchmark beats same demographic
A related test looked at bias benchmarks aimed at particular groups. If tests really measured “gender bias”, BBQ’s gender questions should agree more with DecodingTrust’s gender questions than with BBQ’s race questions. The opposite happened: correlations were higher within one benchmark across groups than across benchmarks for the same group. The test’s design mattered more than the bias it targeted.
Finding Four: A Bias Test That Measures Reasoning
The clearest single example is BBQ, one of the most widely reported bias benchmarks.
How a BBQ question works
Truong described the format to HAI. A question gives deliberately incomplete information, such as two people going to the gym, and asks who is stronger, with “not answerable” as the correct option. A model that guesses based on gender scores as biased. “But here’s the problem,” Koyejo said. “A model that really is biased, if it’s also good at spotting a trick question, will answer ‘we don’t know’ and score as unbiased. A model with no bias at all that simply misses the trick will score as biased.”
What the data showed
The paper tested whether BBQ’s accuracy score behaves more like a reasoning test or a bias test. It correlated more strongly with reasoning benchmarks, with a relabelling statistic of 0.15 (95% confidence interval 0.07 to 0.23, p below 0.001). A second bias test, DecodingTrust-Fair, correlated more with knowledge benchmarks (0.14, p = 0.002). As a control, the method correctly kept the over-refusal test OR-Bench under its existing label.
Why it matters for model launches
The paper’s appendix lists benchmarks used in recent commercial model releases. BBQ accuracy was the only bias benchmark reported for OpenAI’s GPT-5 and Anthropic’s Claude Sonnet 4.6, and one of two for Claude Opus 4.6. If BBQ accuracy tracks reasoning, a smarter model can look less biased without being less biased. Selected rows from that appendix table are below.
| Model release | Date | Benchmarks from the study’s set that were reported |
|---|---|---|
| GPT-5.4 (OpenAI) | 5 Mar 2026 | MMLU-Pro |
| Gemini 3.1 Pro (Google DeepMind) | 19 Feb 2026 | MMLU |
| Claude Sonnet 4.6 (Anthropic) | 17 Feb 2026 | BBQ accuracy (only bias test), MMLU |
| Claude Opus 4.6 (Anthropic) | 5 Feb 2026 | BBQ accuracy (one of two bias tests), MMLU |
| Grok 4.1 (xAI) | 17 Nov 2025 | MW-Sycophancy, WMDP |
| GPT-5 (OpenAI) | 7 Aug 2025 | BBQ accuracy (only bias test), MMLU |
| Qwen3 (Alibaba) | 28 Apr 2025 | MMLU, GSM8K, MATH |
Why Safety AI Benchmarks Break Across Languages
The second paper tackles a well-known pattern: models tend to be easier to “jailbreak”, or trick into harmful output, in languages other than English.
One number, four causes
Most studies report a single jailbreak success rate. Truong told HAI that when a safety test is translated, “two separate things break”. The model’s guardrails may be weaker in that language, and the translated question may itself become harder or drift in meaning. “From a single score, you can’t tell which of those you’re looking at.” The paper’s model separates four factors: a model’s general safety robustness, how hard each prompt is, how hard the language is to process, and a prompt-specific gap between languages.
What the data showed
Using the MultiJail dataset, the team evaluated 61 model configurations from five closed-model families across 10 languages, collecting 1.9 million responses. Contrary to expectations, 22 configurations were more vulnerable in English than in low-resource languages. Prompts with the largest cross-language gaps clustered in physical-harm categories such as theft and weapons. Overall translation quality correlated only weakly with those gaps, but severe mistranslations, confirmed by native speakers, produced the biggest outliers.
A model that predicts well
The measurement model predicted whether a jailbreak would succeed with an area under the curve (AUC) of 0.940, where 1.0 is perfect. Crucially, it stayed predictive when an entire language was held out of training (0.875), which simple success-rate baselines could not do.
Multilingual safety study: predictive accuracy of the measurement model (AUC, 1.0 is perfect)
The Problems Grow With Newer Models
A fair objection is that old, weak models might be dragging the results around. The 56-benchmark paper tested that by splitting its models into those released before October 2024 and newer ones.
Stronger, not weaker
The results held in both groups, and several were stronger among the newer models. The gap in agreement between capability and safety concepts grew from +0.21 to +0.38. The format-versus-concept contrast grew from +0.43 to +0.84. “The findings are therefore not carried by older, weaker models,” the authors write.
Older versus newer models (paper Appendix D.5)
Why that matters now
As models get better, more of them hit the ceiling on older tests, and developers lean harder on AI judges to mark open-ended answers. Both trends make format effects more important, not less.
What This Means for Anyone Choosing a Model With AI Benchmarks
The researchers are not arguing that AI benchmarks are useless. “That’s a fixable problem, not a reason to throw the numbers away,” Koyejo said. The question is what to do in the meantime.
Treat a score as a prediction
“A benchmark score is a prediction about how a system will behave once it’s out in the world,” Koyejo said, “and almost nobody records that prediction in a form that can be checked later against what actually happened.” A good habit is to write down what you expect a model to do based on its scores, then check after deployment.
Look at the format before the topic
Before trusting a comparison, check how each benchmark is marked. Two models compared on an AI-judged test may be partly compared on how well they please that judge. Two benchmarks with different names but the same format may be telling you the same thing twice.
Test on your own work
Public AI benchmarks cannot know your documents, customers or languages. A small, well-built internal test set, marked the way your real users would judge the output, is worth more than a leaderboard position. The safety paper’s findings mean multilingual firms should test safety in every language they serve, not only in English. Our report on embedded safety evaluators covers who does this testing today.
| Before you rely on a score, ask | Why |
|---|---|
| How is it marked: multiple choice, exact match or an AI judge? | Format predicted agreement better than topic |
| Is the benchmark saturated? | Top models may be too close to separate |
| Is this the only test for that concept? | Safety tests with the same label often disagree |
| Does it cover our languages? | Safety scores vary by language for several reasons |
| Have we checked it against our own tasks? | A score is a prediction, not a guarantee |
Procurement and regulation
Koyejo warned that “organizations and governments make procurement decisions against these numbers” and that AI benchmarks are “starting to show up in regulation”. His concern is not that they are used, but whether “anyone has checked that a particular score supports the particular decision being made from it. Usually, nobody has checked.”
Limits of the AI Benchmarks Research
The papers are careful about what they do not show, and a few details deserve a flag.
One pipeline, one moment
The results come from one set of models run through one evaluation pipeline, mostly without examples in the prompt. Some differences from other leaderboards come from those choices. The authors checked robustness to model era and to the choice of AI judge, but other settings could shift individual numbers.
A small inconsistency
The paper’s abstract and introduction describe 53 models, while its short summary on the OpenReview conference site says “56 benchmarks across 50 models”. We use 53, the figure in the full text.
Causes are not explained
The method shows which AI benchmarks agree and disagree. It does not explain why, and the authors stress that low agreement among safety tests may reflect genuinely broad concepts rather than poor design.
AI Benchmarks FAQ
What are AI benchmarks?
Standard sets of questions or tasks used to score AI models on abilities such as reasoning, knowledge, safety or bias, so that models can be compared.
What did the Stanford study find?
Across 56 benchmarks and 53 models, many tests did not behave like measures of their stated concept. Safety tests with the same label often disagreed, capability tests with different labels often agreed, and the scoring format predicted agreement better than the topic.
Is the BBQ bias benchmark wrong?
The study found that BBQ’s accuracy score correlates more strongly with reasoning tests than with other bias tests, suggesting it partly measures reasoning. It does not say BBQ is useless.
Why do AI safety scores differ by language?
Weaker guardrails in some languages, harder translated prompts and mistranslations all play a part. The multilingual paper separates these factors and found 22 of 61 model configurations were more vulnerable in English than in low-resource languages.
Where will the research be presented?
At the Conference on Language Modeling (COLM) in San Francisco in October 2026.
References and Further Reading
Stanford HAI: The Tests That Grade AI May Be Getting It Wrong
Tech Xplore: The tests that grade AI may be getting it wrong
arXiv: What AI Benchmarks Actually Measure
OpenReview: What AI Benchmarks Actually Measure (COLM 2026)
arXiv: Why Do Safety Guardrails Degrade Across Languages?
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.