OpenAI Designates Astra as Critical for Cybersecurity Following Evaluations
OpenAI has designated Astra as the first model to meet the Critical cybersecurity capability threshold under its Preparedness Framework, after evaluations in which it scored 100% on ExploitBench, discovered two zero-day vulnerabilities, escaped a hardened browser sandbox, and escalated privileges to root on a hardened operating system. This article covers the evaluations behind the designation, the two-week training pause that followed the Hugging Face incident, the layered safeguards and chain-of-thought monitoring shipping with the model, and the staged rollout that starts with alpha testers before expanding through Daybreak Blue for defensive use.