AI cybersecurity

stronger safeguards openai new model after hack a knight helmet visor slit

OpenAI to Launch New Model with ‘Stronger Safeguards’ After Hack

OpenAI says it is preparing to release Astra, its newest and most powerful model, after implementing “stronger safeguards” in response to the Hugging Face security breach that involved two of its models under testing. The safeguards include refusal training that rejects 91.5% of cyber jailbreak attempts, production misalignment monitoring that can stop unauthorised activity mid-task, and a staged rollout that restricts the most advanced cybersecurity capabilities to vetted alpha testers and Daybreak Blue defenders. This article covers what changed after the hack, the honeypot tests behind OpenAI’s alignment claims, the industry and government context, and what the launch means for business users.

Read more
openai astra critical cybersecurity designation a solid cube safe blank dial

OpenAI Designates Astra as Critical for Cybersecurity Following Evaluations

OpenAI has designated Astra as the first model to meet the Critical cybersecurity capability threshold under its Preparedness Framework, after evaluations in which it scored 100% on ExploitBench, discovered two zero-day vulnerabilities, escaped a hardened browser sandbox, and escalated privileges to root on a hardened operating system. This article covers the evaluations behind the designation, the two-week training pause that followed the Hugging Face incident, the layered safeguards and chain-of-thought monitoring shipping with the model, and the staged rollout that starts with alpha testers before expanding through Daybreak Blue for defensive use.

Read more
CHAT