CISPA

ai agent manipulation stubborn agents multi agent networks a cartwheel lying flat with a thick hub and eight spokes

When Stubbornness Becomes a Security Problem: How AI Agents Can Manipulate Others

A CISPA Helmholtz Center study borrowed the Friedkin-Johnsen opinion model from sociology, pointed it at six families of large language model, and found that a single stubborn, persuasive agent can hijack an entire multi-agent network. With the attacker at the hub of a star topology, attack success rate reached 1.00 on Gemini-3-Flash and 0.99 on GPT-OSS-120B, against unattacked baselines of 0.02 and 0.06. This article covers the mechanism, the per-model and per-topology numbers, the three structural defences the maths predicts, the trust-adaptive mechanism the team built, and the adaptive attacker that pushes warm-up trust to 0.95.

Read more
CHAT