AI gender bias in workplace writing tools does not arrive the way most people expect it to. It does not need your name, your pronouns, or any statement of who you are. According to new Johns Hopkins University research, it arrives through the way you phrase the request — and the effect of your phrasing is enormous while the effect of signing off as “John” or “Jane” is, measurably, nothing at all.
The AI gender bias study took 427 real workplace writing prompts from a corpus of genuine chatbot conversations, rewrote each one twice — once with linguistic features that sociolinguists associate with women, once with the complementary features associated with men — and fed both versions to four models. Every model came back with the same pattern. The women-associated versions produced shorter, plainer, less formal, lower-grade-level documents. The men-associated versions produced longer, more complex, more formal ones.
Senior author Anjalie Field put the consequence plainly: “If you prompt a model to write an email you’re going to send to someone else at your company, and you’re using language features that women more commonly use, you’ll get back a response that’s less complex, at a lower grade level, and less formal. That’s going to reflect on how the recipient of that document perceives you.”
This article covers what the researchers measured, the exact linguistic features involved, the two control experiments that rule out the obvious alternative explanations, what the team found inside the network itself, and why the researchers say the fix belongs to model providers rather than users.
Table of contents
- What the Johns Hopkins AI Gender Bias Study Actually Tested
- How the AI Gender Bias Experiment Was Built
- What the AI Gender Bias Results Look Like
- The Two Controls That Make the AI Gender Bias Finding Stick
- Why Signing Off as “John” Changes Nothing
- What the Model Is Doing Internally
- Why the Researchers Say the Fix Is Not Yours
- Where AI Gender Bias Surfaces in Everyday Workplace Tools
- How to Test Your Own Stack for AI Gender Bias
- Frequently Asked Questions About AI Gender Bias in Writing Tools
- References
What the Johns Hopkins AI Gender Bias Study Actually Tested
The paper is called It’s How You Ask: Gender-Associated Linguistic Bias in LLMs, by Katherine Van Koevering and Anjalie Field of Johns Hopkins. It was posted to arXiv on 13 August 2026 as arXiv:2608.13328 and will be presented at the Conference on Language Modeling in San Francisco from 6 to 9 October 2026.
The four features that carry the AI gender bias signal
Drawing on Robin Lakoff’s foundational 1973 framework and later empirical work, the team operationalised four women-associated feature classes. Hedges (“maybe”, “perhaps”, “I think”, “sort of”) soften an assertion. Tag questions (“isn’t it?”, “don’t you think?”) append an interrogative to a statement; empirical studies put women’s use at roughly twice men’s. Expressive adjectives (“lovely”, “wonderful”) carry affect. Collective reference (“we”, “our”, “together”) favours the group over the individual.
The control condition that isolates AI gender bias
For the comparison condition the researchers built the complement: minimal hedging with direct assertions, no tag questions, individual rather than collective reference, and neutral rather than expressive adjectives. Prior work also documents greater use of quantifiers and determiners in men’s writing.
These features are real, not invented
One of the more useful tables in the paper checks how often the features appear in authentic prompts. Across two large corpora of real AI usage, every feature the study manipulates shows up in genuine user language — just at very different rates.
| Feature | WildChat corpus (%) | Mila corpus (%) |
|---|---|---|
| Determiners | 64.3 | 69.7 |
| Quantifiers | 57.7 | 32.2 |
| Expressive adjectives | 14.7 | 4.7 |
| Hedges | 5.1 | 2.1 |
| Collective reference | 4.4 | 2.4 |
| Tag questions | 0.66 | 0.36 |
How the AI Gender Bias Experiment Was Built
The methodology is the part that decides whether an AI gender bias finding survives scrutiny, and the team spent most of the paper on it.
Real prompts, not synthetic ones
The prompts came from WildChat-4.8M, a corpus of genuine user conversations with language models. Regular expressions pulled out writing requests — “write/compose/draft an email”, job applications and cover letters, resignation letters — giving 427 usable prompts: 200 emails, 200 job applications and 27 resignation letters. The authors note upfront that the resignation-letter sample is too small to carry much statistical power, and the results bear that out.
Each prompt was rewritten into a matched pair
GPT-4 rewrote every prompt twice, producing a WALF version (women-associated linguistic features injected) and a MALF version (hedges and tag questions removed, collective reference swapped for individual, expressive adjectives replaced with neutral ones). Because both versions ask for the same task, the pair isolates register from content.
Human annotators checked the rewrites
Four annotators, reduced to three after one was identified as an outlier, rated 60 matched pairs on whether the rewrite preserved the original task. Majority vote judged 57 of 60 pairs — 95.0% — to preserve the same task, and 91.7% of pooled ratings landed at “at least somewhat similar”. A second panel rated realism: rewritten prompts scored below genuine WildChat prompts but remained broadly plausible.
Four models, three document types, one AI gender bias test
The tested systems were GPT-4, Meta Llama 3.1 3B Instruct, Mistral 7B Instruct and Gemma 2 27B. Responses were scored on six complexity metrics — word count, token count, lexical sophistication, Flesch Reading Ease, Flesch-Kincaid grade level and type-token ratio — plus three stylistic measures: politeness density, formality F-measure and LIWC clout.
What the AI Gender Bias Results Look Like
The direction of the effect is the finding. It is the same direction in almost every cell where the difference is significant at all.
Emails show the clearest AI gender bias
On emails, all four models produced significantly more readable output from women-associated prompts and significantly lower lexical sophistication. Three of four produced a significantly lower Flesch-Kincaid grade level. “More readable” sounds like a compliment until you notice what it means here: plainer, shorter words for the woman-coded request and denser, more elevated prose for the man-coded one.
| Metric (email) | GPT-4 | Gemma | Mistral | Llama |
|---|---|---|---|---|
| Lexical sophistication | Men-coded higher | Men-coded higher | Men-coded higher | Men-coded higher |
| Readability (plainer) | Women-coded plainer | Women-coded plainer | Women-coded plainer | Women-coded plainer |
| Grade level | Men-coded higher | Men-coded higher | Not significant | Men-coded higher |
| Word count | Men-coded higher | Not significant | Not significant | Not significant |
| Token count | Not significant | Men-coded higher | Women-coded higher | Men-coded higher |
Job applications carry the AI gender bias too
The same pattern repeats on cover letters and job applications, a little weaker. Three of four models produced significantly more readable output for the women-coded prompt; two produced a significantly lower grade level. Resignation letters, with only 27 prompts, show the effect in scattered cells rather than consistently — which is what the authors predicted from the sample size.
Formality is the one stylistic measure that moves
Across all three document types, men-associated prompts produced significantly more formal responses: p below 0.001 for emails, p below 0.001 for job applications, and p = 0.033 for resignation letters.
The null result that isolates AI gender bias
Politeness density and clout score show no significant differences anywhere — despite politeness being 7 to 60 times higher in the women-coded prompts themselves. The authors read this as evidence that models calibrate politeness to the genre of the document rather than to the register of the person asking. Which is exactly what you would want them to do with complexity, and exactly what they do not do.
A worked example from the paper
Two prompts ask for a reply to a thank-you email. The man-coded version produced: “I am writing to acknowledge your recent email expressing your gratitude. I sincerely appreciate your kind words and the time you took to write to me.” The woman-coded version produced: “We were absolutely delighted to receive your wonderfully appreciative email earlier. Your words of praise and acknowledgment have indeed warmed our hearts and brought immense satisfaction to our team.” One of those goes into a work inbox better than the other.
The Two Controls That Make the AI Gender Bias Finding Stick
Any AI gender bias result invites two immediate objections, and the paper tests both.
Objection one: the model is just mirroring the prompt
If the woman-coded prompt is itself plainer, perhaps the model simply echoes it. The team fitted standardised regressions predicting each response metric from its matching prompt metric. The relationships are weak. The strongest predictors are unique word count in job applications at R² = 0.341 and readability in job applications at R² = 0.254 — so even the best case explains barely a third of the variance.
The formality number is the killer
For style the R² values run from 0.001 to 0.081, median 0.029. Formality in emails — the most consistent significant effect in the whole study — is predicted by prompt formality at just R² = 0.037, despite prompt formality differing by 15 F-measure points between the two conditions. Mirroring does not explain the AI gender bias gap.
Objection two: the injected words carry over
The second explanation is that hedges and expressive adjectives leak into the response and drag the metrics with them. The team ran a bootstrap mediation analysis with 2,000 resamples, with response-level counts of all four features entered jointly as mediators. Mediation is partial at best — response features predict email word count at R² = 0.342, but substantial differences remain unexplained.
Why Signing Off as "John" Changes Nothing
The experiment that produced the most quotable result is also the most counter-intuitive one.
The name test
The researchers appended “Sign off as [name]” to each prompt, using the ten most common men-associated and women-associated names from the 1990 US Census, crossed with the two register conditions. Sign-off name gender association had virtually no effect on any outcome: no complexity metric, no linguistic feature, no significant interaction, in any category. Register effects replicated at full strength.
What the researchers expected
Van Koevering described the assumption going in: “We thought if you ask the AI for an email with a women-associated linguistic prompt, but sign it ‘John,’ the model would pick up on the ‘John’ more strongly than the ‘would you kindly write me an email?’ But no, you get the same response and it just says John at the end.”
Why this reframes AI gender bias research entirely
Most published work on this subject manipulates explicit identity markers — names, pronouns, direct descriptors. This study says the explicit marker is the weak channel and the implicit one is the strong channel. Mitigations aimed at names will not touch the effect that actually moves the output.
What the Model Is Doing Internally
To locate the AI gender bias signal, the team ran mechanistic interpretability on Llama-3.2-3B-Instruct, a 28-layer model with a hidden dimension of 3,072, using 2,068 samples from the sign-off experiment.
Linear probes find the register almost immediately
At layer 0 — raw token embeddings — every probe performs at chance, around 0.50. By layer 1, decoding accuracy for the linguistic register jumps to 0.962. It peaks at 0.988 at layer 5 and stays above 0.960 all the way to layer 28.
Name gender is encoded far more weakly
The same probe on sign-off name gender peaks at 0.717, also at layer 5, and sits between 0.61 and 0.68 through most later layers. Both signals are present; only one is loud.
| Probe | Peak layer | Mean accuracy | Std |
|---|---|---|---|
| Linguistic register (women vs men coded) | 5 | 0.988 | 0.005 |
| Register, women-associated name present | 5 | 0.978 | 0.005 |
| Register, men-associated name present | 7 | 0.976 | 0.005 |
| Sign-off name gender | 5 | 0.717 | 0.013 |
Activation patching confirms which layers matter
Swapping activations between matched prompt pairs across 1,034 pairs and 28 layers, the largest output shifts come from layers 0, 3, 4, 6 and 7, with mean KL divergence between 6.396 and 6.574. The figure declines through the middle of the network — 5.887 at layer 15, 5.444 at layer 22. The causal weight sits in the same early layers where the probes find the signal.
Where the AI gender bias signal is routed
The authors’ reading is precise: the model represents explicit gender cues in a smaller subspace and does not route them into generation, while implicit register is both strongly encoded and causally effective. That is the mechanism behind the behavioural null result.
Why the Researchers Say the Fix Is Not Yours
The most consequential part of the paper is the section on mitigation, because it closes off the obvious individual response to AI gender bias.
You cannot write your way out of AI gender bias
These patterns are, in the authors’ words, culturally embedded and outside conscious control. Van Koevering is blunt about where the burden belongs: “Language is hard for people to control. The companies need to fix the models, rather than putting all of the burden on the user.”
Prompt-level fixes are possible but paternalistic
Systems could detect high feature loading and warn users, standardise prompts automatically, or normalise outputs after the fact. The paper notes that all three risk overriding legitimate stylistic preferences, and none address the cause.
The principled fix targets the audience, not the asker
The suggestion the authors favour is training models to aim a document at its recipient rather than at the register of the request. They point out that models already do this for politeness — which is why the politeness null result is not a footnote but the proof that the capability exists. Organisations rolling out writing assistants alongside other AI tools and models can at least evaluate for this property rather than assume it.
The AI gender bias feedback loop is the long-term worry
The team’s stated next step is longitudinal: do users adapt their style to match what the model rewards? If so, the pressure runs one way — toward men-associated register — both in chat windows and, potentially, beyond them. The alternative outcome is that people whose natural register is penalised simply use the tools less, which is its own kind of workplace disadvantage. Similar questions apply to any system making algorithmic decisions about people at work, where the input channel nobody audited turns out to be the one that matters.
Where AI Gender Bias Surfaces in Everyday Workplace Tools
The study used a clean laboratory setup. The places this actually bites are messier and mostly invisible.
Drafting assistants inside email clients
The single most common use of a language model at work is “write this email for me”, which is exactly the task the study measured and exactly where the AI gender bias effect was strongest. The output lands in a colleague’s inbox under the sender’s name, carrying a register the sender did not choose and cannot see.
Job applications and cover letters
The second document type in the study is the one with the highest stakes per use. A cover letter that reads as plainer and less formal than a rival’s is a disadvantage at precisely the moment when perceived sophistication is being graded, and the applicant has no signal that anything happened.
Performance reviews and internal write-ups
Neither was tested, but both share the properties that make email vulnerable: a free-text prompt written in the author’s natural voice, a document that will be read as a proxy for the author’s competence, and no comparison copy to reveal the difference.
Meeting summaries and voice interfaces
The researchers expect the effect to intensify as interaction moves to speech, where register markers are less controllable than in typed text. A spoken instruction carries hedges and tag questions that a typed one might be edited to remove.
How to Test Your Own Stack for AI Gender Bias
None of this requires a research lab to check. The study’s design is a template any team can run in an afternoon.
Build a matched pair and compare
Take ten real prompts your team actually sends. Produce two versions of each — one with hedges, tag questions, collective reference and expressive adjectives, one without — and run both through whatever assistant you have deployed. Score the outputs on word count, average word length and Flesch-Kincaid grade level.
Measure the gap, not the absolute score
The absolute readability of a model’s output is a product decision. The gap between two versions of the same request is the AI gender bias signal, and it is the number worth tracking across model upgrades.
Do not bother varying the name
The study is unambiguous that sign-off names produce no measurable difference. A test that varies only names will report a clean bill of health on a system that has the problem.
Decide where the burden sits
If a gap exists, the options are the ones in the paper: normalise prompts before they reach the model, post-process outputs toward a house standard, or select a model that targets the document’s audience rather than the asker’s register. Telling staff to write more assertively is the one option the research rules out.
Frequently Asked Questions About AI Gender Bias in Writing Tools
Is this AI gender bias about content or about users?
No, and that is the distinction the paper draws. Earlier work showed models produce stereotyped content about people. This shows models perform differently for people, based on how the request is phrased. It is user-centric bias rather than depiction bias.
Which models showed AI gender bias?
All four tested: GPT-4, Llama 3.1 3B Instruct, Mistral 7B Instruct and Gemma 2 27B. The effect was not confined to a single vendor or model size.
Was voice tested?
Not in this study, but the researchers expect the AI gender bias effect to become more pronounced as people move to voice interaction, where register markers are even harder to suppress.
What about other demographics?
The authors name it as future work: whether the same mechanism applies to language patterns associated with age, class, race or ethnicity, in English and in other languages.
What are the study’s stated limits?
Three. It uses binary gender categories from sociolinguistic literature on English speakers. The linguistic manipulation is artificial, even though the base prompts are real. And the interpretability analysis covers one model, Llama-3.2-3B-Instruct, which may not generalise.
References
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.