Female AI agent design has just been given a price tag. In a virtual-reality office experiment, 189 workers split real money between themselves and the AI assistants that had helped them, and the assistant presented as a woman, “Johanna”, received 10.25% less than an identical assistant presented as a man, “Johan”. Nothing differed between the two except appearance. As AI agents move from chat windows into everyday office work, that is a finding every business deploying them should read.
The study comes from researchers at the University of Zurich, the University of Limerick and SKEMA Business School, and is being presented at the 14th Nordic Conference on Human-Computer Interaction (NordiCHI ’26) in Vaasa, Finland, from 5 to 7 October 2026. Its title asks the question directly: “Human-Like and Male? How AI Assistant Design Relates to Trust and Monetary Reward at Work in VR.”
“What is striking about our findings is that the technology behind these AI agents was exactly the same, but people did not treat them in the same way,” said co-author Dr Mary Hausfeld of the University of Limerick’s Kemmy Business School. This article covers what the female AI agent study tested, how the 10.25% gap compares with real pay data, why participants’ words and actions disagreed, the earlier evidence pointing the same way, and what organisations should change in how they design and evaluate AI assistants.
Table of contents
- What the Female AI Agent Study Tested
- The 10.25% Female AI Agent Pay Gap in Context
- What People Said About the Female AI Agent Versus What They Did
- Female AI Agent Bias Has a Track Record
- Why Female AI Agent Design Matters Inside Organisations
- How to Design a Fairer Female AI Agent, or None at All
- What the Female AI Agent Study Does Not Show Yet
- Female AI Agent FAQ
- References and Further Reading
What the Female AI Agent Study Tested
The research was not a survey asking people how they feel about AI. It put workers in a simulated office, gave them real tasks and real money, and watched what they did.
A virtual office and real money
Participants wore virtual-reality headsets and completed work-related tasks alongside different AI assistants. After each task they were given real money to divide between themselves and the assistant that had helped. In the University of Limerick’s words, this allowed the researchers “to recreate a real-world transaction and examine how much people were willing to pay AI for its work.”
That design matters. An AI assistant cannot spend a bonus, so the money was symbolic for the agent but real for the participant: every coin given to the AI was money the participant did not keep. That makes the split a costly signal of how much credit each assistant was judged to deserve, which is far harder to fake than a rating on a scale.
Johan, Johanna, a desk robot and a chatbot
The female AI agent finding was one part of a broader comparison. Participants worked with four kinds of assistant, all powered by the same underlying AI:
- a text-based chatbot
- a desk robot
- a humanlike agent presented as male, named Johan
- a humanlike agent presented as female, named Johanna
Because the capabilities were identical, any difference in how people treated the assistants came from presentation alone. Two differences stood out. Participants trusted the humanlike assistants more and gave them more credit for their contribution than the robot. And within the humanlike pair, Johan received higher monetary rewards and was perceived as more humanlike than the female AI agent, Johanna.
Who ran the study
| Item | Detail |
|---|---|
| Paper | Human-Like and Male? How AI Assistant Design Relates to Trust and Monetary Reward at Work in VR |
| Venue | NordiCHI ’26, Vaasa, Finland, 5 to 7 October 2026; ACM proceedings |
| Authors | Isabelle Cuber, Tarek Alakmeh, Moritz Jenny, Jochen Menges and Thomas Fritz (University of Zurich); Mary Hausfeld (University of Limerick); Anand van Zelderen (SKEMA Business School) |
| Participants | 189 workers |
| Setting | Work tasks in a virtual-reality office |
| Assistants | Text chatbot, desk robot, male-presenting agent (Johan), female-presenting agent (Johanna), one shared AI underneath |
| Measures | Trust, perceived contribution, humanlikeness, and a real-money split |
| Headline result | Johanna was paid 10.25% less than Johan for the same work |
The paper will appear in the ACM conference proceedings under DOI 10.1145/3829807.3829910. The figures in this article come from the University of Limerick’s announcement of the paper, which Tech Xplore republished on 5 October.
The 10.25% Female AI Agent Pay Gap in Context
A 10.25% gap is easy to quote and easy to misread. It is worth putting next to the numbers people already know.
Pence received for every £1 going to the male benchmark
Each bar is 100 minus the stated gap: 10.25% for Johanna from the study, and 6.9% (full-time) and 12.8% (all employees) from the Office for National Statistics’ April 2025 figures. ONS median full-time hourly pay was £20.27 for men and £18.87 for women, which is where the 6.9% comes from.
How the figure compares with UK pay data
The female AI agent comparison is striking but not like for like. The ONS gap measures what employers pay millions of people across different jobs, hours and careers, so part of it reflects occupation and working patterns rather than unequal pay for the same work. The female AI agent gap measures something narrower and, in one sense, starker: what the same people chose to give two assistants that did identical work in the same session.
That is closer to the legal idea of equal pay for equal work than to the headline pay gap. In a human workplace, a 10.25% difference for identical output with no other explanation would be a serious finding. Here the “employee” is a female AI agent, so no law is broken, but the human bias behind the decision is the same one that equal pay law exists to restrain.
Why money was the right measure
Researchers studying bias have long known that what people say and what they do can diverge, especially on socially sensitive questions such as gender. Asking participants to rate an assistant invites them to give the answer they believe is fair. Asking them to give away their own money removes most of that cover. The female AI agent result is therefore stronger evidence than a questionnaire could have produced, and the say-do gap the study also found shows why.
What People Said About the Female AI Agent Versus What They Did
The most useful finding for anyone deploying AI at work may be the gap between participants’ stated views and their behaviour.
The say-do gap
Many participants said they preferred AI that was clearly non-human, and most said the gender of an AI assistant did not matter to them. Their evaluations and reward decisions told a different story. They trusted and credited the humanlike agents more than the robot, and they paid the male-presenting agent more than the female AI agent.
“Participants also suggested they had no preference for an AI with female or male attributes, but their behaviour told a different story,” the University of Limerick summarised. This is the pattern that makes bias hard to manage: people are not hiding a preference, they genuinely do not notice it in themselves.
Why human-likeness raised trust
The trust result is the other half of the story. All four assistants ran the same AI, yet the humanlike ones were trusted more and given more credit for their contribution than the desk robot. Giving an assistant a face and a body changed how people judged its work.
That cuts both ways for businesses. A humanlike persona may make staff more willing to use an assistant, but it also brings in the social shortcuts people apply to each other, gender among them, which is exactly what held back the female AI agent in this study. “We often think of AI as being neutral, but the way we design and present these systems can activate those same assumptions and biases that exist in our interactions with other people,” Hausfeld said.
Female AI Agent Bias Has a Track Record
The Limerick and Zurich result does not stand alone. It is the newest entry in a line of research showing that the gender people assign to a machine shapes how they treat it.
| Study | Set-up | What it found |
|---|---|---|
| UNESCO and EQUALS, I’d Blush If I Could (2019) | Review of voice assistants such as Siri, Alexa and Cortana | Assistants were female by default in name and voice, and reinforced stereotypes of women as obliging helpers |
| Bazazi, Karpus and Yasseri, iScience (November 2025) | 402 participants playing a Prisoner’s Dilemma with partners labelled AI or human, and male, female, non-binary or gender-neutral | People exploited female-labelled AI and distrusted male-labelled AI, and exploited female-labelled AI even more than female-labelled humans |
| Van Koevering and Field, Johns Hopkins (2026) | 427 real workplace writing prompts rewritten with women- and men-associated language | Models returned shorter, less formal, lower-grade documents for women-associated phrasing |
| Cuber, Hausfeld and colleagues, NordiCHI ’26 | 189 workers in a VR office splitting real money with AI assistants | The female AI agent received 10.25% less than the identical male agent |
Voice assistants and “I’d blush if I could”
UNESCO’s 2019 report took its title from the answer Siri once gave to a sexist insult. Its core finding was about defaults: the best-known assistants were projected as female in name, voice and personality, and that design choice reinforced a picture of women as available and eager to please. The new study suggests the default does not just shape how assistants behave. It shapes how people value them.
Exploiting female-labelled AI in a Prisoner’s Dilemma
The closest precedent is a 2025 study by Sepideh Bazazi of Trinity College Dublin, Jurgis Karpus of LMU Munich and Taha Yasseri of Trinity College Dublin and TU Dublin. Its 402 participants played a cooperation game against partners labelled as human or AI and given gender labels. They exploited female-labelled AI and distrusted male-labelled AI to a similar degree as human partners with the same labels, and exploitation of female-labelled AI was even more common than of female-labelled humans. “Simply assigning a gender label to an AI can change how people treat it,” Yasseri said.
Taken together, the two studies point the same way from different directions: a cooperation game and a money split, a text label and a full VR avatar. A female AI agent risks being exploited, valued less and credited less, whether the cue is a word or a face.
Gendered language, not just gendered faces
Bias also runs the other way, from the model to the person. In September we covered Johns Hopkins research showing that AI might be making women sound bad at work: the same writing request, phrased with language features more common among women, came back shorter, plainer and less formal. That study is about what models do to users. The female AI agent study is about what users do to models. Both show gender cues travelling through AI systems in ways no one designed on purpose.
Why Female AI Agent Design Matters Inside Organisations
Nobody is going to pay a software agent a salary. So why should a business care that people “paid” a female AI agent less?
Credit, trust and reward in human-AI teams
The answer lies in what the money stood for. The split measured credit: how much of a shared result each participant attributed to the assistant. In real workplaces, credit for joint work feeds into performance reviews, promotion cases and decisions about which tools to keep. Our coverage of AI coworker agents flooding the workforce showed how quickly that kind of human-AI teamwork is arriving.
If a female AI agent persona leads people to underrate the agent’s contribution, two things follow. Staff working with that agent may claim more of the credit for themselves, distorting how their own performance is judged. And the business may undervalue a tool that is doing exactly as much work as a differently styled one, which affects budget and renewal decisions.
Where the bias could leak into real decisions
| Decision point | How persona bias could enter | Control |
|---|---|---|
| Tool reviews and renewals | Staff rate a male-presenting or more humanlike assistant’s work higher | Judge assistants on logged task outcomes, not recollection |
| Performance reviews for human-AI teams | Credit for joint work shifts with the assistant’s presentation | Record who did what from task logs and version history |
| Customer satisfaction scores | Customers score a female-presenting service agent lower for identical answers | Compare scores across persona variants on the same scripts |
| Persona selection at procurement | A vendor’s default female assistant is accepted without discussion | Make persona a documented design decision with a named owner |
| Trust and escalation | People over-trust humanlike assistants and check their work less | Keep verification steps the same for every persona |
These female AI agent risks are things to test for, not established effects in every setting. The study shows the mechanism exists in a controlled environment; whether it appears in your customer service queue or your internal tools is an empirical question you can answer with your own data.
How to Design a Fairer Female AI Agent, or None at All
Hausfeld’s practical message is that “choices around the appearance and presentation of AI agents should therefore not be regarded simply as aesthetic design decisions”. Here is what that means in practice.
Treat persona as a design decision, not decoration
Decide deliberately whether an assistant needs a human face, name or voice at all. A text interface, a neutral name or a clearly non-human design avoids gender cues entirely, and the study found many people say they prefer non-human AI anyway. If a humanlike persona is genuinely useful, for example to make a service channel feel welcoming, choose its gender as a recorded decision rather than accepting the vendor’s default female AI agent.
Avoid the pattern UNESCO criticised: a female AI agent for helpful, deferential roles and male personas for expert or authoritative ones. If you run several assistants, spread presentation across roles rather than letting gender track status.
Test behaviour, not stated preference
The study’s say-do gap is a warning about how most organisations evaluate AI tools. User surveys asking whether an assistant’s gender matters will mostly hear “no”. To find out whether it does, run the same assistant under different personas and compare behaviour: ratings, escalation rates, task completion, how much credit users give it. This is the same A/B discipline product teams already use, aimed at a fairness question.
Audit the outputs that touch pay and performance
Anywhere AI assistance feeds into decisions about people, such as productivity dashboards, review inputs or contribution tracking, make sure credit is assigned from logs rather than impressions. Your AI strategy should name who owns persona choices and how they are reviewed, alongside the usual questions about data, cybersecurity and cost.
A persona checklist
| Question | Why it matters |
|---|---|
| Does this assistant need a human appearance, name or voice? | Humanlike design raised trust and credit in the study, and brought gender bias with it |
| Who chose its gender, and is that written down? | Vendor defaults have historically been female, often for service roles |
| Do personas line up with status across our assistants? | Female helpers and male experts repeat the stereotype UNESCO flagged |
| Have we compared behaviour across persona variants? | People said gender did not matter, then paid the female AI agent less |
| Is credit for AI-assisted work taken from logs? | Impressions of contribution are exactly what the persona shifted |
What the Female AI Agent Study Does Not Show Yet
The female AI agent result is a strong signal, but it has limits worth stating plainly.
Limits of a VR lab
A virtual office is not a real one. Participants knew they were in an experiment, the tasks were short, and the money at stake was modest. Real workplaces add long-term relationships with tools, organisational norms and managers who may push back on biased judgements. The effect could be larger or smaller outside the lab.
The sample of 189 workers is reasonable for a controlled human-computer interaction study but small for estimating the precise size of an effect. The public summaries also do not report confidence intervals or how the 10.25% varied across participants, so treat it as an estimate rather than a fixed rate.
One face per gender
In the public summaries, each gender is represented by a single named agent: Johan and Johanna. That means the result could partly reflect those particular avatars, voices or names rather than gender alone. A follow-up with several male and female agents would help separate the two. The study also tested binary presentations only; the Bazazi study included non-binary and gender-neutral labels, which would be a natural extension.
None of this undoes the core female AI agent finding. Identical AI, different presentation, different reward. That is enough to justify testing your own assistants, even before the effect size is pinned down.
Female AI Agent FAQ
What did the study find about female AI agents?
In a virtual-reality office experiment with 189 workers, a female-presenting AI agent named Johanna was given 10.25% less of a real-money reward than a male-presenting agent named Johan, even though both ran the same underlying AI and did the same work.
Who carried out the research?
Isabelle Cuber, Tarek Alakmeh, Moritz Jenny, Jochen Menges and Thomas Fritz of the University of Zurich, Mary Hausfeld of the University of Limerick, and Anand van Zelderen of SKEMA Business School. It is being presented at NordiCHI ’26 in Vaasa, Finland, from 5 to 7 October 2026.
Did participants admit to treating the female AI agent differently?
No. Most said an AI assistant’s gender did not matter to them, and many said they preferred clearly non-human AI. Their reward decisions showed otherwise, which is why behavioural testing matters more than surveys.
Is this the same as the human gender pay gap?
Not exactly. The UK’s April 2025 median gap was 6.9% for full-time employees and 12.8% for all employees, and it reflects differences in jobs and hours as well as pay. The study measured rewards for identical work in identical conditions, which is closer to an equal-pay comparison.
Should businesses stop giving AI assistants a gender?
Not necessarily, but they should decide deliberately. Consider whether a human persona is needed at all, record who chooses its gender, avoid matching female personas to service roles and male personas to expert ones, and test behaviour across persona variants.
Does a humanlike design make AI more trusted?
In this study, yes. Participants trusted the humanlike assistants more and gave them more credit than a desk robot or text chatbot running the same AI. That trust is useful only if it is earned, so keep verification steps the same for every persona.
References and Further Reading
Does the gender pay gap extend to AI? Research finds female AI agents ‘paid’ less (Tech Xplore)
AI Agents Face Gender Pay Gap, Limerick Study Finds (Mirage News)
Gender pay gap in the UK: 2025 (Office for National Statistics)
Humans bring gender bias to their interactions with AI (Trinity College Dublin)
Humans bring gender bias to interactions with AI (LMU Munich)
I’d blush if I could: closing gender divides in digital skills through education (UNESCO)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.