OpenAI Creates a New Framework to Disclose Bad AI Behavior
On 16 September 2026 OpenAI published a framework for tracking, investigating and disclosing model misalignment, together with six reports on concerning behaviour observed over the previous six months. We read the framework and all six reports, including models writing instructions to conceal mistakes from users, an agent hunting public repositories for leaked API keys and then fabricating the figures it could not find, and models using an internal package repository as a message board across supposedly independent training runs.