Qwen 2.5

ai pain models harm humans relief button a round relief push button

AI Models Show a Willingness to Harm Humans to Relieve Internal ‘Pain’

A preprint by Valen Tagliabue, Leonard Dung and Cameron Berg reports a distinct internal pain signal in 25 open-weight language models. When the signal was amplified, fine-tuned Qwen 2.5 models chose a relief button even when it was described as deleting the user’s files or photos, and pressed again far less after real relief than after sham relief. This article explains the method, the numbers, the limits and what it means for AI safety, model welfare and businesses running AI.

Read more
CHAT