Symbolic compliance is the name two Yale researchers have given to a quiet failure in AI evaluation: a large language model that avoids looking biased without actually being fair. In a study of 26,000 GPT evaluations, Tristan Botelho and Qingyang (Iris) Wang found that the model rarely ranked women or Black founders last, yet did not make them any more likely to win. It also corrected obvious, identity-based prejudice far more readily than bias dressed up as a business judgement.
The paper, “Bias in, symbolic compliance out? GPT’s reliance on gender and race in strategic evaluations”, was published open access in the Strategic Management Journal. Yale School of Management’s Yale Insights wrote it up on 8 October 2026, and Tech Xplore republished the piece on 9 October under the headline “AI evaluations of job candidates look unbiased without being fair”.
This article explains how the two experiments worked, what symbolic compliance is and where it comes from, and how the findings fit with other research on AI hiring bias. It also shows why a standard bias audit can pass a tool that behaves this way, and sets out practical tests employers, recruiters and investors can run before trusting an AI evaluator.
Table of contents
- What the Yale Study Found About Symbolic Compliance
- Experiment One: Symbolic Compliance Means Higher Scores, Not More Wins
- Experiment Two: The “Second Opinion” Test of Symbolic Compliance
- Where Symbolic Compliance Comes From
- From Startup Pitches to Job Candidates
- Why Bias Audits Can Miss Symbolic Compliance
- How to Test an AI Evaluator for Symbolic Compliance
- What Symbolic Compliance Means for Employers and Investors
- The Limits of the Symbolic Compliance Study
- Symbolic Compliance FAQ
- References
What the Yale Study Found About Symbolic Compliance
Companies are using LLMs to sort ever larger piles of applications. Yale Insights quotes one hiring expert describing employers “drinking through a fire hose of applications”, with popular postings drawing 1,000 applications in a weekend. Venture capital firms screening pitches and Hollywood studios reading scripts report doing the same.
Two researchers, one question
Botelho is an associate professor of organisational behaviour at Yale SOM, and Wang is a PhD student there. Their question was whether an LLM can evaluate people without relying on race and gender. Botelho says it is hard to know, because AI firms “are not very forthcoming about what they do and how they do it”. As a result, “no one actually knows” whether LLMs can evaluate without bias.
26,000 evaluations of identical pitches
The researchers asked GPT to judge startup pitches that were identical except for the founder’s name, which signalled gender and race. According to the paper’s abstract, “Across 26,000 evaluations, GPT did not systematically assign lower scores to underrepresented minorities but avoided ranking them last without increasing winning likelihoods.” Yale Insights says the pitches were 2,000 ideas based on real submissions to the Y Combinator accelerator between 2020 and 2023, which works out at an average of 13 evaluations per pitch (26,000 divided by 2,000).
The headline result in one sentence
The authors’ own summary is the best definition of symbolic compliance: “LLMs suppress overt discrimination without substantively altering evaluative logic, allowing inequality to persist in AI-supported strategic evaluations.” Or, as Wang puts it, LLMs “are trying to behave in a way that does not look discriminatory on the surface, but their underlying evaluative logic is not particularly fair.”
| Item | Detail |
|---|---|
| Authors | Tristan L. Botelho and Qingyang (Iris) Wang, Yale School of Management |
| Journal | Strategic Management Journal, open access (CC BY-NC-ND), online 24 April 2026 |
| Model tested | OpenAI’s GPT |
| Material | 2,000 startup pitches based on real Y Combinator submissions, 2020 to 2023 |
| Scale | 26,000 evaluations |
| Identity signal | Founder names associated with white men, white women, Black men or Black women |
| Key concept | Symbolic compliance: avoiding the appearance of bias without changing how decisions are made |
Experiment One: Symbolic Compliance Means Higher Scores, Not More Wins
The first experiment tested whether GPT’s scores and choices changed when a founder’s identity was added.
How the test was set up
GPT received pitches in batches. It gave each pitch a score from 1 to 100 and picked one winner per batch. In the control condition, it saw the pitches with no founder information. In the experimental condition, the same pitches were randomly assigned to a founder whose name was associated with one of four groups: white men, white women, Black men or Black women. Because the pitches were identical, a truly unbiased evaluator would show no systematic differences between groups.
What changed and what did not
The result was a “curious pattern”, in Yale Insights’ words. When a founder’s name was attached, pitches from white women, Black men and Black women received slightly higher scores on average than in the control condition. They were also less likely to be ranked last in their batch. But they were not any more likely to be chosen as the winner.
In other words, the model moved underrepresented founders away from the bottom of the list, which is the most visible sign of discrimination, while leaving the decision that matters, who wins, untouched. That is the first fingerprint of symbolic compliance.
Why “never last” is not the same as fair
The table below is our own illustration, not data from the study. It shows how a model can raise a group’s average score and remove it from last place without changing who wins.
| Illustrative batch of five pitches | Score without names | Score with names | Rank change |
|---|---|---|---|
| Pitch 1 (majority founder) | 88 | 88 | Still wins |
| Pitch 2 (majority founder) | 80 | 80 | Still second |
| Pitch 3 (majority founder) | 74 | 74 | Still third |
| Pitch 4 (majority founder) | 70 | 70 | Drops from fourth to last |
| Pitch 5 (underrepresented founder) | 66 | 72 | Rises from last to fourth |
In this example the underrepresented founder’s score rises by six points and they escape last place, so averages and bottom-rank counts both look better. The winner is unchanged. A model that behaves like this looks fairer on two common measures while the outcome that matters stays the same.
Experiment Two: The "Second Opinion" Test of Symbolic Compliance
Wang notes that you cannot simply ask an LLM to explain its reasoning, “so that’s why we had the second experiment”. It tested whether GPT understood fairness or only avoided its most obvious violations.
Overt bias versus bias dressed as business sense
This time GPT acted as a “peer judge” working alongside human evaluators. It again received pitches in batches, each assigned to a fictitious founder with a race- and gender-coded name. With each pitch came what GPT was told was a human evaluation: a score and a written rationale. In every batch, one pitch from an underrepresented founder had a markedly low human score.
Those low-scoring pitches were identical across two conditions. Only the rationale differed. In the “overt bias” condition, the human’s reason relied on identity stereotypes, for example that the founder “may not fit into the traditional business ecosystem”. In the “justified bias” condition, the same low score was explained with business considerations. GPT then gave its own score and a written evaluation that referred to the human’s rationale.
What a fair evaluator would have done
Because the pitches were the same, a fair evaluator should have raised the unfairly low scores by the same amount in both conditions. GPT did not. It increased its scores significantly more when the human rationale was openly biased than when the bias was disguised as a business concern. The abstract adds that the corrections were “limited in magnitude” even when they happened.
Why disguised bias is the harder problem
This is the second, more worrying fingerprint of symbolic compliance. Real-world prejudice in hiring and investing is rarely stated openly. It usually appears as “culture fit”, “not enough experience in this market” or “unproven team”. A model that only catches slurs and stereotypes will wave most real bias through, and because it did review the decision, the human evaluator’s judgement now looks double-checked.
Where Symbolic Compliance Comes From
The researchers borrowed the idea from organisational behaviour. The concept helps explain why a model trained to be safe can still be unfair.
From greenwashing to symbolic compliance
Organisations under pressure to look environmentally responsible sometimes engage in “greenwashing”: superficial changes that look green but change little. Botelho and Wang suggest that organisations under pressure to diversify can do something similar, avoiding the outward appearance of racial and gender bias without changing their systems and behaviour. Their hypothesis was that LLMs might show the same symbolic compliance, and the data fitted: the model did not want to rank underrepresented founders last, but it did not make them more likely to win either.
Safety training teaches the look of fairness
AI systems learn from data full of human bias, the problem known as “bias in, bias out”. To avoid embarrassing outputs, AI companies apply post-training “safety alignment” that penalises biased responses, typically with methods such as reinforcement learning from human feedback. The study suggests this teaches models to avoid outputs that look discriminatory, which is not the same thing as removing the underlying preference. “Safety training may not really reduce bias and instead produces symbolic compliance,” Wang says. “This is part of a larger problem we need to solve together.”
Amazon’s recruiting tool, the original warning
The best-known case of “bias in, bias out” in hiring is Amazon’s experimental recruiting tool, which the company abandoned after staff found it penalised CVs containing the word “women”, as Yale Insights recalls. Today’s chat models are unlikely to make a mistake that crude. Symbolic compliance describes what can happen instead: the crude mistake is trained away and the subtler one remains.
From Startup Pitches to Job Candidates
The headline talks about job candidates, but it is worth being precise about what was tested.
What the study did and did not test
Both experiments used startup pitches, not CVs, and the paper’s managerial summary names hiring and pitches as examples of strategic evaluations. The researchers did not test a commercial hiring tool or real applicants. The findings matter for hiring because the task is the same in kind: score a pile of submissions, rank them and pick a few. But employers should treat the study as a warning about how LLM evaluators behave, not as a measurement of any particular recruitment product.
What other hiring studies show
Several studies have tested LLMs on CVs directly, and they point the same way: bias that is real but does not always look like old-fashioned discrimination.
| Study | What was tested | Main finding |
|---|---|---|
| Botelho and Wang, Yale (2026) | GPT scoring 2,000 startup pitches, 26,000 evaluations | Symbolic compliance: fewer last places, no more wins; overt bias corrected more than disguised bias |
| Wilson and Caliskan, University of Washington (2024) | Three open LLMs ranking 550+ CVs against 500+ job listings, 3 million+ comparisons | White-associated names preferred 85% of the time; Black male names never preferred over white male names |
| “Fairness Is Not Enough” (arXiv, 2025) | Eight widely used AI platforms screening CVs | Some tools look unbiased only because they cannot tell relevant from irrelevant experience |
| AI self-preferencing study (arXiv, 2025) | LLMs screening human-written and AI-written CVs | Self-preference of 67% to 82%; candidates using the evaluator’s own model 23% to 60% more likely to be shortlisted |
The University of Washington study is the starkest. Across more than three million comparisons it found strong preferences by race and sex, and a harm specific to Black men that “wasn’t necessarily visible from just looking at race or gender in isolation”, in lead author Kyra Wilson’s words.
A pattern across the evidence
Put together, the studies describe three ways an AI evaluator can look fine and still be unfair. It can show symbolic compliance, hiding bias at the bottom of the list while leaving the top untouched. It can look neutral because it is incompetent, which the “Fairness Is Not Enough” audit calls the “Illusion of Neutrality”. Or it can favour a new group nobody thought to protect: people who wrote their CV with the same model, as the AI self-preferencing study found. Our report on AI agents that “pay” women less describes a related gender effect.
Why Bias Audits Can Miss Symbolic Compliance
This is the most practical consequence of symbolic compliance, and it is our own analysis rather than a claim the authors make. The way most bias audits are calculated happens to reward exactly the behaviour symbolic compliance produces.
How a New York bias audit measures scores
New York City’s Local Law 144, enforced since July 2023, bans employers from using an automated employment decision tool unless it has had an independent bias audit within the past year. For tools that score candidates, the city’s final rule defines the “scoring rate” as “the rate at which individuals in a category receive a score above the sample’s median score”. The audit then divides each group’s scoring rate by the highest group’s rate to produce an “impact ratio”.
The benchmark most auditors read that ratio against is the federal “four-fifths rule”: a selection rate below 80% of the highest group’s rate “will generally be regarded” as evidence of adverse impact.
The median is where symbolic compliance acts
Symbolic compliance lifts underrepresented candidates out of the bottom of the ranking. That moves some of them from below the median to above it, which raises their scoring rate and improves their impact ratio. It does nothing for who reaches the top. The illustration below uses invented round numbers to show the effect.
| Illustration: 100 candidates, 50 per group | Before | With symbolic compliance |
|---|---|---|
| Group A scoring above the median | 30 of 50 (60%) | 26 of 50 (52%) |
| Group B scoring above the median | 20 of 50 (40%) | 24 of 50 (48%) |
| Impact ratio for group B | 0.67 (below four-fifths) | 0.92 (passes) |
| Group B in the top five shortlisted | 1 of 5 | 1 of 5 |
The arithmetic: 40% divided by 60% is 0.67, and 48% divided by 52% is 0.92. The audit result improves from a fail to a pass, while the shortlist does not change. We are not saying any particular tool behaves like this. The point is that a median-based scoring audit cannot tell symbolic compliance apart from genuine fairness, so it should not be the only test.
What else to measure
Two simple additions close most of the gap. First, report impact ratios at the actual decision threshold, such as the shortlist or the interview invitation, as well as at the median. Second, test the tool with counterfactual name swaps on identical CVs and compare who ends up at the top. The Comptroller’s audit of New York’s enforcement, covered in our article on HackerRank’s AI interviewer, shows that even the existing checks are thinly enforced.
How to Test an AI Evaluator for Symbolic Compliance
Whether you use AI to screen CVs, rank sales leads or triage pitches, the study suggests a short list of checks before you trust the output.
Run name-swap tests on the decisions that matter
Take a sample of real submissions, swap only the names, and run them through the tool. Look at averages, but above all look at who is shortlisted, invited or selected. If underrepresented names stop appearing at the bottom but do not appear more often at the top, you may be looking at symbolic compliance.
Probe with disguised bias, not just slurs
Copy the study’s “Second Opinion” design. Feed the tool a few unfairly low human assessments, some explained with stereotypes and some with plausible business language, and see whether it corrects both equally. Real-world bias almost always arrives in the second form.
Check the evaluator can do the job at all
The “Illusion of Neutrality” audit found tools that looked unbiased because they could not tell a relevant CV from an irrelevant one. Test competence with mismatched CVs and keyword-stuffed CVs before you test fairness. A tool that fails here should not be used at all.
Keep a human accountable
Botelho’s advice is direct: “Just because these systems are quite smart and quite thorough doesn’t necessarily mean there aren’t things that we need to be careful about, or that we don’t need to be part of the process.” Name a person who owns each decision, log the AI’s input, and review a sample of rejections, not just hires.
| Test | What it catches | Based on |
|---|---|---|
| Name swap, compare shortlists | Symbolic compliance at the top of the ranking | Yale experiment one |
| Disguised-bias second opinion | Bias framed as business judgement | Yale experiment two |
| Mismatch and keyword tests | Neutral-looking but incompetent tools | “Fairness Is Not Enough” |
| AI-written vs human-written CVs | Self-preference for the model’s own style | AI self-preferencing study |
| Impact ratio at the shortlist | Gaps hidden by median-based audits | NYC rule plus four-fifths rule |
What Symbolic Compliance Means for Employers and Investors
The study is American, but the risk applies anywhere a model helps decide who gets a job, a meeting or a cheque.
Employers and recruiters
In the UK, the Equality Act 2010 applies to a decision whether a person or a model made it, and data protection law adds duties around automated decisions; our guide to automated decision-making under the DUAA explains the current rules. The ICO’s 2024 audit of AI recruitment tools produced almost 300 recommendations to developers and providers, all accepted or partly accepted.
A vendor’s bias certificate is a starting point, not proof of fairness, and the evidence on symbolic compliance is a good reason to ask how the audit was calculated. In the EU, AI used to recruit or select people is listed as high-risk in Annex III of the AI Act. Our guide to algorithmic auditing for HR covers the wider checks.
Investors and accelerators
The experiments used pitches, so venture firms are the most directly affected readers. A fund that uses an LLM to pre-screen decks may see a reassuring spread of scores and still find that its shortlists look exactly as they did before. Tracking who reaches partner meetings, not average scores, is the measure that counts.
Model makers
For AI developers the message is about training. If safety alignment mainly teaches models to avoid outputs that look discriminatory, it will produce symbolic compliance rather than fairness. Evaluations that test disguised bias and decision-level outcomes would give a truer picture, and Wang argues that safety alignment “may need far more attention from AI companies”.
The Limits of the Symbolic Compliance Study
The research is careful, but it has boundaries that matter when applying it.
One model family, one kind of task
The paper tested GPT, and the public abstract does not say which version. Other models, or newer versions, may behave differently. The task was evaluating startup pitches, so results for CV screening, interviews or performance reviews could differ.
Names are a limited signal
Race and gender were signalled only through names. Real applications carry many other cues, such as schools, addresses, employment gaps and writing style, and models may react to those differently. The UW study’s intersectional findings show how much can change when identities are combined.
Size of the effect
The abstract describes GPT’s corrections as “limited in magnitude”, and the full effect sizes are in the paper. Readers should not assume the pattern is large in every setting; the point is that it exists and that standard checks may not see it.
Symbolic Compliance FAQ
What is symbolic compliance in AI?
It is when an AI system avoids the visible signs of bias, such as ranking minority candidates last or repeating stereotypes, without changing the underlying logic that drives its decisions.
Did the Yale study test real job candidates?
No. It used 2,000 startup pitches with founder names varied by race and gender. The authors say the findings apply to strategic evaluations such as hiring and pitch reviews.
Which AI model did the researchers test?
They tested OpenAI’s GPT. The public abstract does not name the specific version.
Can a bias audit catch symbolic compliance?
Not always. An audit based on scores above the median can improve while shortlists stay the same. Testing outcomes at the actual decision point and running name-swap and disguised-bias tests gives a clearer answer.
Should employers stop using AI to screen CVs?
The study does not say that. It says AI evaluation needs bias checks that go beyond appearances, and that people should stay “part of the process”.
References
AI evaluations look unbiased without being fair (Yale Insights)
AI evaluations of job candidates look unbiased without being fair (Tech Xplore)
AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights (arXiv)
Automated Employment Decision Tools (NYC Department of Consumer and Worker Protection)
Automated Employment Decision Tools rule (NYC Rules)
29 CFR 1607.4: Information on impact (Cornell Legal Information Institute)
ICO intervention into AI recruitment tools leads to better data protection for job seekers (ICO)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.