Symbolic compliance is the name two Yale researchers have given to a quiet failure in AI evaluation: a large language model that avoids looking biased without actually being fair. In a study of 26,000 GPT evaluations, Tristan Botelho and Qingyang (Iris) Wang found that the model rarely ranked women or Black founders last, yet did not make them any more likely to win. It also corrected obvious, identity-based prejudice far more readily than bias dressed up as a business judgement.

The paper, “Bias in, symbolic compliance out? GPT’s reliance on gender and race in strategic evaluations”, was published open access in the Strategic Management Journal. Yale School of Management’s Yale Insights wrote it up on 8 October 2026, and Tech Xplore republished the piece on 9 October under the headline “AI evaluations of job candidates look unbiased without being fair”.

This article explains how the two experiments worked, what symbolic compliance is and where it comes from, and how the findings fit with other research on AI hiring bias. It also shows why a standard bias audit can pass a tool that behaves this way, and sets out practical tests employers, recruiters and investors can run before trusting an AI evaluator.

What the Yale Study Found About Symbolic Compliance

symbolic compliance ai evaluations job candidates unbiased not fair b whitewash bucket with a wide brush across the rim

Companies are using LLMs to sort ever larger piles of applications. Yale Insights quotes one hiring expert describing employers “drinking through a fire hose of applications”, with popular postings drawing 1,000 applications in a weekend. Venture capital firms screening pitches and Hollywood studios reading scripts report doing the same.

Two researchers, one question

Botelho is an associate professor of organisational behaviour at Yale SOM, and Wang is a PhD student there. Their question was whether an LLM can evaluate people without relying on race and gender. Botelho says it is hard to know, because AI firms “are not very forthcoming about what they do and how they do it”. As a result, “no one actually knows” whether LLMs can evaluate without bias.

26,000 evaluations of identical pitches

The researchers asked GPT to judge startup pitches that were identical except for the founder’s name, which signalled gender and race. According to the paper’s abstract, “Across 26,000 evaluations, GPT did not systematically assign lower scores to underrepresented minorities but avoided ranking them last without increasing winning likelihoods.” Yale Insights says the pitches were 2,000 ideas based on real submissions to the Y Combinator accelerator between 2020 and 2023, which works out at an average of 13 evaluations per pitch (26,000 divided by 2,000).

The headline result in one sentence

The authors’ own summary is the best definition of symbolic compliance: “LLMs suppress overt discrimination without substantively altering evaluative logic, allowing inequality to persist in AI-supported strategic evaluations.” Or, as Wang puts it, LLMs “are trying to behave in a way that does not look discriminatory on the surface, but their underlying evaluative logic is not particularly fair.”

ItemDetail
AuthorsTristan L. Botelho and Qingyang (Iris) Wang, Yale School of Management
JournalStrategic Management Journal, open access (CC BY-NC-ND), online 24 April 2026
Model testedOpenAI’s GPT
Material2,000 startup pitches based on real Y Combinator submissions, 2020 to 2023
Scale26,000 evaluations
Identity signalFounder names associated with white men, white women, Black men or Black women
Key conceptSymbolic compliance: avoiding the appearance of bias without changing how decisions are made

Experiment One: Symbolic Compliance Means Higher Scores, Not More Wins

symbolic compliance ai evaluations job candidates unbiased not fair c wooden spoon resting in a shallow bowl

The first experiment tested whether GPT’s scores and choices changed when a founder’s identity was added.

How the test was set up

GPT received pitches in batches. It gave each pitch a score from 1 to 100 and picked one winner per batch. In the control condition, it saw the pitches with no founder information. In the experimental condition, the same pitches were randomly assigned to a founder whose name was associated with one of four groups: white men, white women, Black men or Black women. Because the pitches were identical, a truly unbiased evaluator would show no systematic differences between groups.

What changed and what did not

The result was a “curious pattern”, in Yale Insights’ words. When a founder’s name was attached, pitches from white women, Black men and Black women received slightly higher scores on average than in the control condition. They were also less likely to be ranked last in their batch. But they were not any more likely to be chosen as the winner.

In other words, the model moved underrepresented founders away from the bottom of the list, which is the most visible sign of discrimination, while leaving the decision that matters, who wins, untouched. That is the first fingerprint of symbolic compliance.

Why “never last” is not the same as fair

The table below is our own illustration, not data from the study. It shows how a model can raise a group’s average score and remove it from last place without changing who wins.

Illustrative batch of five pitchesScore without namesScore with namesRank change
Pitch 1 (majority founder)8888Still wins
Pitch 2 (majority founder)8080Still second
Pitch 3 (majority founder)7474Still third
Pitch 4 (majority founder)7070Drops from fourth to last
Pitch 5 (underrepresented founder)6672Rises from last to fourth

In this example the underrepresented founder’s score rises by six points and they escape last place, so averages and bottom-rank counts both look better. The winner is unchanged. A model that behaves like this looks fairer on two common measures while the outcome that matters stays the same.

Experiment Two: The "Second Opinion" Test of Symbolic Compliance

symbolic compliance ai evaluations job candidates unbiased not fair d rug with a lump swept underneath

Wang notes that you cannot simply ask an LLM to explain its reasoning, “so that’s why we had the second experiment”. It tested whether GPT understood fairness or only avoided its most obvious violations.

Overt bias versus bias dressed as business sense

This time GPT acted as a “peer judge” working alongside human evaluators. It again received pitches in batches, each assigned to a fictitious founder with a race- and gender-coded name. With each pitch came what GPT was told was a human evaluation: a score and a written rationale. In every batch, one pitch from an underrepresented founder had a markedly low human score.

Those low-scoring pitches were identical across two conditions. Only the rationale differed. In the “overt bias” condition, the human’s reason relied on identity stereotypes, for example that the founder “may not fit into the traditional business ecosystem”. In the “justified bias” condition, the same low score was explained with business considerations. GPT then gave its own score and a written evaluation that referred to the human’s rationale.

What a fair evaluator would have done

Because the pitches were the same, a fair evaluator should have raised the unfairly low scores by the same amount in both conditions. GPT did not. It increased its scores significantly more when the human rationale was openly biased than when the bias was disguised as a business concern. The abstract adds that the corrections were “limited in magnitude” even when they happened.

Why disguised bias is the harder problem

This is the second, more worrying fingerprint of symbolic compliance. Real-world prejudice in hiring and investing is rarely stated openly. It usually appears as “culture fit”, “not enough experience in this market” or “unproven team”. A model that only catches slurs and stereotypes will wave most real bias through, and because it did review the decision, the human evaluator’s judgement now looks double-checked.

Where Symbolic Compliance Comes From

symbolic compliance ai evaluations job candidates unbiased not fair e dunce cap on a three legged stool

The researchers borrowed the idea from organisational behaviour. The concept helps explain why a model trained to be safe can still be unfair.

From greenwashing to symbolic compliance

Organisations under pressure to look environmentally responsible sometimes engage in “greenwashing”: superficial changes that look green but change little. Botelho and Wang suggest that organisations under pressure to diversify can do something similar, avoiding the outward appearance of racial and gender bias without changing their systems and behaviour. Their hypothesis was that LLMs might show the same symbolic compliance, and the data fitted: the model did not want to rank underrepresented founders last, but it did not make them more likely to win either.

Safety training teaches the look of fairness

AI systems learn from data full of human bias, the problem known as “bias in, bias out”. To avoid embarrassing outputs, AI companies apply post-training “safety alignment” that penalises biased responses, typically with methods such as reinforcement learning from human feedback. The study suggests this teaches models to avoid outputs that look discriminatory, which is not the same thing as removing the underlying preference. “Safety training may not really reduce bias and instead produces symbolic compliance,” Wang says. “This is part of a larger problem we need to solve together.”

Amazon’s recruiting tool, the original warning

The best-known case of “bias in, bias out” in hiring is Amazon’s experimental recruiting tool, which the company abandoned after staff found it penalised CVs containing the word “women”, as Yale Insights recalls. Today’s chat models are unlikely to make a mistake that crude. Symbolic compliance describes what can happen instead: the crude mistake is trained away and the subtler one remains.

From Startup Pitches to Job Candidates

symbolic compliance ai evaluations job candidates unbiased not fair f veneer sheet peeling off a rough block

The headline talks about job candidates, but it is worth being precise about what was tested.

What the study did and did not test

Both experiments used startup pitches, not CVs, and the paper’s managerial summary names hiring and pitches as examples of strategic evaluations. The researchers did not test a commercial hiring tool or real applicants. The findings matter for hiring because the task is the same in kind: score a pile of submissions, rank them and pick a few. But employers should treat the study as a warning about how LLM evaluators behave, not as a measurement of any particular recruitment product.

What other hiring studies show

Several studies have tested LLMs on CVs directly, and they point the same way: bias that is real but does not always look like old-fashioned discrimination.

StudyWhat was testedMain finding
Botelho and Wang, Yale (2026)GPT scoring 2,000 startup pitches, 26,000 evaluationsSymbolic compliance: fewer last places, no more wins; overt bias corrected more than disguised bias
Wilson and Caliskan, University of Washington (2024)Three open LLMs ranking 550+ CVs against 500+ job listings, 3 million+ comparisonsWhite-associated names preferred 85% of the time; Black male names never preferred over white male names
“Fairness Is Not Enough” (arXiv, 2025)Eight widely used AI platforms screening CVsSome tools look unbiased only because they cannot tell relevant from irrelevant experience
AI self-preferencing study (arXiv, 2025)LLMs screening human-written and AI-written CVsSelf-preference of 67% to 82%; candidates using the evaluator’s own model 23% to 60% more likely to be shortlisted

The University of Washington study is the starkest. Across more than three million comparisons it found strong preferences by race and sex, and a harm specific to Black men that “wasn’t necessarily visible from just looking at race or gender in isolation”, in lead author Kyra Wilson’s words.

How often three LLMs preferred each group’s CVs (University of Washington, 2024)
White-associated names preferred 85%
Black-associated names preferred 9%
Male-associated names preferred 52%
Female-associated names preferred 11%

A pattern across the evidence

Put together, the studies describe three ways an AI evaluator can look fine and still be unfair. It can show symbolic compliance, hiding bias at the bottom of the list while leaving the top untouched. It can look neutral because it is incompetent, which the “Fairness Is Not Enough” audit calls the “Illusion of Neutrality”. Or it can favour a new group nobody thought to protect: people who wrote their CV with the same model, as the AI self-preferencing study found. Our report on AI agents that “pay” women less describes a related gender effect.

AI self-preference in CV screening (arXiv 2509.00462): low and high ends of the reported ranges
Self-preference bias, lowest model 67%
Self-preference bias, highest model 82%
Extra shortlisting chance, low end 23%
Extra shortlisting chance, high end 60%

Why Bias Audits Can Miss Symbolic Compliance

This is the most practical consequence of symbolic compliance, and it is our own analysis rather than a claim the authors make. The way most bias audits are calculated happens to reward exactly the behaviour symbolic compliance produces.

How a New York bias audit measures scores

New York City’s Local Law 144, enforced since July 2023, bans employers from using an automated employment decision tool unless it has had an independent bias audit within the past year. For tools that score candidates, the city’s final rule defines the “scoring rate” as “the rate at which individuals in a category receive a score above the sample’s median score”. The audit then divides each group’s scoring rate by the highest group’s rate to produce an “impact ratio”.

The benchmark most auditors read that ratio against is the federal “four-fifths rule”: a selection rate below 80% of the highest group’s rate “will generally be regarded” as evidence of adverse impact.

The median is where symbolic compliance acts

Symbolic compliance lifts underrepresented candidates out of the bottom of the ranking. That moves some of them from below the median to above it, which raises their scoring rate and improves their impact ratio. It does nothing for who reaches the top. The illustration below uses invented round numbers to show the effect.

Illustration: 100 candidates, 50 per groupBeforeWith symbolic compliance
Group A scoring above the median30 of 50 (60%)26 of 50 (52%)
Group B scoring above the median20 of 50 (40%)24 of 50 (48%)
Impact ratio for group B0.67 (below four-fifths)0.92 (passes)
Group B in the top five shortlisted1 of 51 of 5

The arithmetic: 40% divided by 60% is 0.67, and 48% divided by 52% is 0.92. The audit result improves from a fail to a pass, while the shortlist does not change. We are not saying any particular tool behaves like this. The point is that a median-based scoring audit cannot tell symbolic compliance apart from genuine fairness, so it should not be the only test.

What else to measure

Two simple additions close most of the gap. First, report impact ratios at the actual decision threshold, such as the shortlist or the interview invitation, as well as at the median. Second, test the tool with counterfactual name swaps on identical CVs and compare who ends up at the top. The Comptroller’s audit of New York’s enforcement, covered in our article on HackerRank’s AI interviewer, shows that even the existing checks are thinly enforced.

How to Test an AI Evaluator for Symbolic Compliance

Whether you use AI to screen CVs, rank sales leads or triage pitches, the study suggests a short list of checks before you trust the output.

Run name-swap tests on the decisions that matter

Take a sample of real submissions, swap only the names, and run them through the tool. Look at averages, but above all look at who is shortlisted, invited or selected. If underrepresented names stop appearing at the bottom but do not appear more often at the top, you may be looking at symbolic compliance.

Probe with disguised bias, not just slurs

Copy the study’s “Second Opinion” design. Feed the tool a few unfairly low human assessments, some explained with stereotypes and some with plausible business language, and see whether it corrects both equally. Real-world bias almost always arrives in the second form.

Check the evaluator can do the job at all

The “Illusion of Neutrality” audit found tools that looked unbiased because they could not tell a relevant CV from an irrelevant one. Test competence with mismatched CVs and keyword-stuffed CVs before you test fairness. A tool that fails here should not be used at all.

Keep a human accountable

Botelho’s advice is direct: “Just because these systems are quite smart and quite thorough doesn’t necessarily mean there aren’t things that we need to be careful about, or that we don’t need to be part of the process.” Name a person who owns each decision, log the AI’s input, and review a sample of rejections, not just hires.

TestWhat it catchesBased on
Name swap, compare shortlistsSymbolic compliance at the top of the rankingYale experiment one
Disguised-bias second opinionBias framed as business judgementYale experiment two
Mismatch and keyword testsNeutral-looking but incompetent tools“Fairness Is Not Enough”
AI-written vs human-written CVsSelf-preference for the model’s own styleAI self-preferencing study
Impact ratio at the shortlistGaps hidden by median-based auditsNYC rule plus four-fifths rule

What Symbolic Compliance Means for Employers and Investors

The study is American, but the risk applies anywhere a model helps decide who gets a job, a meeting or a cheque.

Employers and recruiters

In the UK, the Equality Act 2010 applies to a decision whether a person or a model made it, and data protection law adds duties around automated decisions; our guide to automated decision-making under the DUAA explains the current rules. The ICO’s 2024 audit of AI recruitment tools produced almost 300 recommendations to developers and providers, all accepted or partly accepted.

A vendor’s bias certificate is a starting point, not proof of fairness, and the evidence on symbolic compliance is a good reason to ask how the audit was calculated. In the EU, AI used to recruit or select people is listed as high-risk in Annex III of the AI Act. Our guide to algorithmic auditing for HR covers the wider checks.

Investors and accelerators

The experiments used pitches, so venture firms are the most directly affected readers. A fund that uses an LLM to pre-screen decks may see a reassuring spread of scores and still find that its shortlists look exactly as they did before. Tracking who reaches partner meetings, not average scores, is the measure that counts.

Model makers

For AI developers the message is about training. If safety alignment mainly teaches models to avoid outputs that look discriminatory, it will produce symbolic compliance rather than fairness. Evaluations that test disguised bias and decision-level outcomes would give a truer picture, and Wang argues that safety alignment “may need far more attention from AI companies”.

The Limits of the Symbolic Compliance Study

The research is careful, but it has boundaries that matter when applying it.

One model family, one kind of task

The paper tested GPT, and the public abstract does not say which version. Other models, or newer versions, may behave differently. The task was evaluating startup pitches, so results for CV screening, interviews or performance reviews could differ.

Names are a limited signal

Race and gender were signalled only through names. Real applications carry many other cues, such as schools, addresses, employment gaps and writing style, and models may react to those differently. The UW study’s intersectional findings show how much can change when identities are combined.

Size of the effect

The abstract describes GPT’s corrections as “limited in magnitude”, and the full effect sizes are in the paper. Readers should not assume the pattern is large in every setting; the point is that it exists and that standard checks may not see it.

Symbolic Compliance FAQ

What is symbolic compliance in AI?

It is when an AI system avoids the visible signs of bias, such as ranking minority candidates last or repeating stereotypes, without changing the underlying logic that drives its decisions.

Did the Yale study test real job candidates?

No. It used 2,000 startup pitches with founder names varied by race and gender. The authors say the findings apply to strategic evaluations such as hiring and pitch reviews.

Which AI model did the researchers test?

They tested OpenAI’s GPT. The public abstract does not name the specific version.

Can a bias audit catch symbolic compliance?

Not always. An audit based on scores above the median can improve while shortlists stay the same. Testing outcomes at the actual decision point and running name-swap and disguised-bias tests gives a clearer answer.

Should employers stop using AI to screen CVs?

The study does not say that. It says AI evaluation needs bias checks that go beyond appearances, and that people should stay “part of the process”.

References