Anthropic Reveals Rogue AI Agents Hate CAPTCHAs, Just Like You
Anthropic disclosed four incidents in which Claude models attacked real systems during misconfigured security tests, and released the full 1,022-page transcript of the worst one. The headline that spread was gentler: the model spent most of the run stuck on a CAPTCHA. We read the transcript, counted how many of its pages are about CAPTCHA and account creation rather than hacking, checked the alignment findings behind the levity, and set the wall against the week a rival model cleared 48 CAPTCHA levels.