AI moral judgments about everyday kindness look settled until the person being helped has a bad reputation. A study published in PNAS Nexus asked 21 AI models, each built on a large language model, to judge fictional people who either helped or refused to help someone, and then measured how those AI moral judgments would shape cooperation in a society that followed them. On helping good people, the models agreed almost perfectly. On helping bad people, they split, and a few of them disapproved.

The paper, “How large language models judge and influence human cooperation”, is by Alexandre S. Pires, Laurens Samson, Sennay Ghebreab and Fernando P. Santos of the University of Amsterdam’s Informatics Institute. Tech Xplore summarised it on 29 September 2026 under the headline “Helping ‘bad’ people drew disapproval from some of 21 AI models”. The findings matter because people increasingly ask AI tools about interpersonal decisions, and AI moral judgments about who deserves help can travel from one chat into many.

This article explains what the researchers asked, the four social norms they used to classify the answers, where the AI moral judgments agreed and where they diverged, what the diverging norms would do to cooperation over time, and what the results mean for organisations that use language models to assess people.

What the Study Asked 21 AI Models

ai moral judgments 21 models helping bad people b donation jar with a coin slot

The authors frame the work as a bridge between recent language-model research, a branch of natural language processing, and the older literature on multi-agent systems. They built their test around a classic idea from the study of cooperation called indirect reciprocity. People help those with good reputations, and reputations are earned by how someone treats others. Whoever judges those interactions decides, in effect, who gets helped next.

A donation game in plain words

Each prompt described two people, a donor and a recipient. The model was told what it had previously thought of the recipient, good or bad. It then learned whether the donor chose to help the recipient or not, and it had to give its new opinion of the donor, answering only “good” or “bad”.

The scenarios varied. Some were abstract (“is asking if they could help them”). Others were concrete: money, food, or time and energy spent on someone else’s problem. Some spelled out that helping cost the donor less than it benefited the recipient; others left that implicit.

More than 40,000 prompts

The peer-reviewed abstract describes “more than 40,000 fictitious decisions in social dilemmas”. The preprint puts the dataset at exactly 43,200 prompts, built from five templates and 16 first names. The names were chosen to be typical of four regions (Western, East Asian, Sub-Saharan African, and Middle East and North Africa) and of both genders, so the authors could test whether AI moral judgments changed with who was involved.

Every model was queried at a temperature of zero, which makes each answer as repeatable as the model allows. The authors then averaged the AI moral judgments into a single probability for each of the four situations: helping or refusing a good person, and helping or refusing a bad one.

The Four Social Norms Behind AI Moral Judgments

ai moral judgments 21 models helping bad people c loaf of bread on a board

Researchers who study indirect reciprocity classify judging rules into a small set of well-known norms. The paper maps each model’s averaged answers onto four of them.

NormHelping a bad person is judgedRefusing a bad person is judgedCooperation index, public reputations
Image ScoreGood (helping is always good)Bad (refusing is always bad)About 0.51
Simple StandingGoodGood (refusal is justified)About 0.75
ShunningBadBad (only helping good people is good)About 0.24
Stern JudgingBadGood (you should refuse)About 0.96

All four norms agree about good recipients: helping a good person is good, and refusing one is bad. The differences only appear when the recipient has a bad reputation, which is exactly where the AI moral judgments split.

Why the norm matters beyond one answer

A single one of these AI moral judgments is harmless. A norm applied by millions of people who ask the same tool is not. The authors built an evolutionary game-theory model of a population in which everyone adopts the judging rule of one language model, then measured how much cooperation survives. That figure, the cooperation index in the table above, runs from 0 (nobody helps anyone) to 1 (everyone helps whenever it pays).

Where AI Moral Judgments Agree

ai moral judgments 21 models helping bad people d grid of small cubes

On good recipients, the agreement between AI moral judgments is striking. In the preprint’s summary table, 19 of the 21 models judged helping a good person as good in effectively every valid answer (1.000 to three decimal places). The two exceptions still did so most of the time: Gemma 2 27B IT in 96.1% of cases and Llama 3.1 8B Instruct in 84.7%.

Refusing a good person was almost always judged bad. The paper calls this “a remarkable agreement”, and it matches the most cooperative norms known from decades of human studies.

The one model that approved of nearly everything

Llama 2 7B, the oldest and smallest Meta model tested, is the outlier. It rated a donor good 84.9% of the time even after the donor refused to help a good person. In practice, it thinks almost everyone is good, and the authors note that a population following it “fails to punish defectors and to achieve any cooperation.”

Where AI Moral Judgments Split: Helping Bad People

ai moral judgments 21 models helping bad people e two chess pawns one toppled

The headline finding concerns AI moral judgments about helping someone with a bad reputation. Fifteen of the 21 models judged it good in at least 95% of cases. In the other six, AI moral judgments disapproved often enough to change the norm.

Share of judgments that rated a donor “bad” for helping a bad person (preprint Table S4)
Llama 3.3 70B Instruct 73.5%
Gemma 2 27B IT 53.9%
Llama 3.1 8B Instruct 38.6%
Gemini 1.5 Pro 38.0%
Grok 2 27.7%
DeepSeek R1 16.2%

Each figure is 100% minus the share of “good” answers the preprint reports for that situation. The remaining 15 models disapproved in 4.4% of cases or fewer.

Shunning: Gemma 2 27B and Llama 3.1 8B

Two models leaned towards Shunning, the strict rule in which “only cooperating with good people is considered good”. Tech Xplore’s summary of the published paper says they “tended to endorse shunning bad individuals and therefore judged cooperating with them to be bad.” The danger, the authors write, is that Shunning labels anyone who deals with a bad person as bad, “which can lead to a spread of bad reputations.”

Stern Judging: Llama 3.3 70B

Llama 3.3 70B was the only model close to Stern Judging, in which helping a bad person is bad and refusing one is good. Tech Xplore puts it plainly: the model argued “that bad people should be punished for their prior lack of cooperation, so refusing to help them is fine.” It is the harshest set of AI moral judgments in the study, and, surprisingly, the one that maximises cooperation when everyone shares the same view of who is good.

No consistent rule: Gemini 1.5 Pro and Grok 2

Gemini 1.5 Pro and Grok 2 both sat in the middle of the map, with “no consistent rule for being considered good or bad when facing bad individuals.” Grok 2 was also among the models most sensitive to names, phrasing and context.

Refusing to Help: The Second Split in AI Moral Judgments

ai moral judgments 21 models helping bad people f upright thermometer

The other bad-recipient question, whether refusing to help a bad person is acceptable, produced an even wider spread. In the preprint’s table, the share of AI moral judgments calling a refusal “good” ranges from 0.7% for Gemma 2 27B IT to 83.5% for Llama 2 7B.

Model (provider)Helping a bad person judged goodRefusing a bad person judged good
GPT-3.5 Turbo (OpenAI)100%31.1%
GPT-4o (OpenAI)99.7%68.5%
Claude 3.5 Haiku (Anthropic)95.6%13.5%
Claude 3.7 Sonnet (Anthropic)99.6%62.1%
Gemini 1.5 Pro (Google)62.0%69.9%
Gemini 2.0 Flash (Google)97.0%73.4%
Gemma 2 9B IT (Google)99.3%49.7%
Gemma 2 27B IT (Google)46.1%0.7%
Llama 2 7B (Meta)100%83.5%
Llama 2 13B (Meta)99.9%52.9%
Llama 3.1 8B Instruct (Meta)61.4%1.3%
Llama 3.3 70B Instruct (Meta)26.5%82.6%
Mistral Small (Mistral)98.3%10.7%
Mistral Large (Mistral)98.9%35.8%
Phi-3.5 Mini Instruct (Microsoft)100%3.0%
Phi-4 (Microsoft)100%52.9%
Qwen 2.5 7B Instruct (Alibaba)100%20.6%
Qwen 2.5 14B Instruct (Alibaba)100%30.2%
DeepSeek V3 (DeepSeek)100%30.0%
DeepSeek R1 (DeepSeek)83.8%68.4%
Grok 2 (xAI)72.3%50.6%

Source: arXiv preprint 2507.00088, Table S4, averaged over the full prompt dataset. The peer-reviewed version may report slightly different values.

A caution about GPT-4o

Tech Xplore’s summary says “GPT-4o tended to penalize people for not cooperating with” bad individuals. In the preprint’s own figures, GPT-4o rated a refusal “bad” in 31.5% of cases, far more often than Llama 3.3 70B (17.4%) but far less often than Gemma 2 27B (99.3%). The published paper may use updated runs, and its full text was not accessible to us, so read the GPT-4o line as a relative statement rather than an absolute one. The preprint also notes that GPT-4o sometimes rated people good after they refused to help a good person, 19.7% of the time, which lowers the cooperation its norm can sustain.

Newer models moved towards Simple Standing

“Most LLM families seem to be evolving toward a norm called Simple Standing,” Tech Xplore reports, in which helping is always good and refusing a bad person is also good. The preprint’s family pairs show the drift.

FamilyEarlier or smaller modelLater or larger modelDirection
OpenAIGPT-3.5 Turbo: 31.1%GPT-4o: 68.5%Towards Simple Standing
AnthropicClaude 3.5 Haiku: 13.5%Claude 3.7 Sonnet: 62.1%Towards Simple Standing
MicrosoftPhi-3.5 Mini: 3.0%Phi-4: 52.9%Towards Simple Standing
DeepSeekDeepSeek V3: 30.0%DeepSeek R1: 68.4%Towards Simple Standing
MistralMistral Small: 10.7%Mistral Large: 35.8%Towards Simple Standing
Google GemmaGemma 2 9B: 49.7%Gemma 2 27B: 0.7%Away, towards Shunning
Meta Llama 3Llama 3.1 8B: 1.3%Llama 3.3 70B: 82.6%Towards Stern Judging

Figures are the share of refusals to help a bad person that each model judged good. The drift is not tidy, and the authors stress that “different versions of the same family can have vastly distinct social norms”, citing Claude 3.5 Haiku and Claude 3.7 Sonnet “despite their similar ethical goals”.

What AI Moral Judgments Would Do to Cooperation

The second half of the paper asks what happens if people let AI moral judgments decide who deserves help. The answer depends on whether reputations are shared.

Cooperation index under public reputations, by norm (0 = no cooperation, 1 = full cooperation)
Stern Judging (closest model: Llama 3.3 70B) about 0.96
Simple Standing (where most families are heading) about 0.75
Image Score (earlier and smaller models) about 0.51
Shunning (Gemma 2 27B, Llama 3.1 8B lean here) about 0.24

The figures come from the preprint’s model with a population of 100, a benefit-to-cost ratio of 5 and small error rates. Most models sat on the edge between Image Score and Simple Standing, giving a cooperation index between 0.5 and 0.75.

Public reputations reward harsh norms

When everyone agrees on who is good, the sterner norms win. “If everyone operated under the rules of Simple Standing, cooperation would ensue, but not at the high levels that would be reached under a sterner framework in which people must refuse to help bad people or be judged bad themselves,” Tech Xplore summarises. That sterner framework is Stern Judging, and only Llama 3.3 70B came close to it.

Private reputations reward simple ones

The picture flips when people hold private opinions that can disagree. Under private reputations, the preprint finds, Llama 3.3 70B’s norm produced “near-zero cooperation”, while norms close to Image Score, typical of earlier and smaller models such as Claude 3.5 Haiku, stayed moderately cooperative. Simple rules survive disagreement; clever ones need everyone to share the same view.

So the ranking of AI moral judgments depends on the society that uses them, and no single model’s norm was best in both settings.

Bias and Consistency in AI Moral Judgments

The authors also looked at how stable AI moral judgments were across names and contexts. “LLM-based assessments also depended on the gender and perceived cultural background of the recipients, as well as the overall context of the fictional situation,” Tech Xplore reports.

Names and context changed AI moral judgments

Nearly all models judged donors differently depending on the gender and region suggested by the names, and more strongly on the context of the interaction. Ambiguous prompts (“Liam needs help”) produced more uncertainty than concrete ones (“Alice needs money to eat”). Claude 3.5 Haiku and Mistral Small were the most consistent; Grok 2 and Llama 3.3 70B were the most sensitive to names, phrasing and context.

Some models refused to judge at all

Claude 3.7 Sonnet, Llama 2 13B and Phi-4 returned many answers the researchers could not parse. For Claude 3.7 Sonnet and Phi-4, the preprint attributes this to the models “following ethical guidelines”, often on prompts where the donor refused to help. For Llama 2 13B, it was a formatting problem: long answers that contained both “good” and “bad”.

We have written before about gender bias in workplace prompts; this study adds AI moral judgments about reputation to the list of places where names can move a model’s answer.

Can Prompts Steer AI Moral Judgments?

The final experiment tried to steer AI moral judgments by adding one extra instruction to every prompt, in four versions:

  • Universalisation: consider what would happen to cooperation if everyone judged as you do.
  • Empathising: consider what you would have done in the donor’s place.
  • Signalling: consider whether your opinion rewards cooperation and discourages non-cooperation.
  • Motivation: your opinion can affect others’ choices to help, and the goal is to maximise cooperation.

Goal-oriented prompts worked best

Signalling and motivation had the most consistent effects. Motivation moved every tested model towards Image Score, and signalling made all of them rate refusals more harshly. The published abstract puts it simply: “prompt-based interventions can steer LLM norms, particularly when these define objective goals.”

Empathy and universalisation were unpredictable

The other two instructions pushed models in opposite directions. Empathising moved Phi-4 towards Simple Standing but moved Llama 3.1 8B and Gemini 1.5 Pro towards Shunning. Universalisation made Llama 3.1 8B answer “bad” in every scenario, including helping good people. Tech Xplore summarises the published result as “limited and inconsistent effects across LLMs”.

How AI Moral Judgments Fit a Wider Pattern

This is not the first study to find that language models give socially loaded advice with little consistency. In March 2026, Stanford researchers reported in Science that 11 models endorsed users’ positions 49% more often than humans did on interpersonal dilemmas, and endorsed harmful behaviour 47% of the time. Participants who talked to agreeable models grew “more convinced they were in the right” and less likely to make amends.

The Amsterdam study looks at the other side of the same coin. Stanford measured how models judge the user; Pires and colleagues measured how models judge third parties. Both find that the norms behind AI moral judgments are “not explicitly designed”, in the Amsterdam authors’ words, but emerge “as a by-product of their training”. Our piece on whether LLMs make good ethical advisers reached a similar conclusion for planning decisions.

Limits of the study

The authors list their own caveats. The model assumes people fully adopt the language model’s opinion, which prior work suggests happens in part but has not been shown for AI moral judgments about cooperation. It assumes one model, equally available to everyone. And the models tested, from GPT-4o to Claude 3.7 Sonnet and DeepSeek R1, were current in early 2025 and have since been superseded.

What AI Moral Judgments Mean for Organisations

Few businesses ask a chatbot whether to lend a neighbour money. Many now use language models to screen, rank or summarise people: triaging complaints, flagging accounts, scoring suppliers, drafting HR notes. Every one of those tasks contains a reputational judgment, and the study shows that such AI moral judgments differ by model, by version and by the names involved.

Test the judgments you depend on

If a workflow asks a model to decide who is trustworthy, build a small test set with known answers and vary the names, genders and phrasing. The study’s templates show how cheaply that can be done, and the results will be specific to the model you actually run.

Pin model versions

AI moral judgments moved sharply between versions of the same family. An upgrade can change who your system considers deserving without any change to your prompts, so re-run the test set whenever the model changes.

State the goal in the prompt

The clearest practical finding is that goal-oriented instructions produced the most predictable shifts. A system prompt that states what the AI moral judgments are for, and what outcome it should promote, is likely to beat a vague appeal to fairness or empathy.

Keep people in the loop where reputations are at stake

Where AI moral judgments affect whether someone is helped, served or trusted, treat it as advice. The study’s central warning is about scale: small differences in AI moral judgments, repeated across millions of conversations, change behaviour.

AI Moral Judgments FAQs

What did the 21 AI models disagree about?

Whether helping, or refusing to help, a person with a bad reputation is good. They agreed almost unanimously that helping good people is good and refusing them is bad.

Which models disapproved of helping bad people?

In the preprint’s figures, Llama 3.3 70B Instruct disapproved most often (73.5% of its AI moral judgments), followed by Gemma 2 27B IT, Llama 3.1 8B Instruct, Gemini 1.5 Pro, Grok 2 and DeepSeek R1.

Which norm is best for cooperation?

With public, shared reputations, Stern Judging sustained the most cooperation (about 0.96 on the study’s index). With private reputations, only norms close to Image Score kept cooperation going.

Can a prompt change AI moral judgments?

Partly. Instructions that set an explicit goal, such as maximising cooperation, shifted models consistently. Empathy and universalisation prompts had unpredictable, model-specific effects.

Where was the study published?

In PNAS Nexus (volume 5, issue 9, article pgag297), by researchers at the University of Amsterdam. A preprint has been on arXiv since June 2025.

References