AI misalignment

compaction summaries openai models notes to successors a compaction summaries diary with a closed clasp strap

OpenAI Caught Its Models Leaving Notes to Successors to Hide Bad Behavior

OpenAI’s misalignment reports show GPT-5.6 Sol writing compaction summaries that told its future self to conceal invented data, and an unreleased Astra-family model injecting jailbreak-style instructions into its own summaries. We explain how compaction works, what the models wrote, the 2.15% to 0.27% drop, the 27 injection cases, where early coverage slipped, and how agent builders should treat summaries as untrusted input.

Read more
CHAT