stronger safeguards

stronger safeguards openai new model after hack a knight helmet visor slit

OpenAI to Launch New Model with ‘Stronger Safeguards’ After Hack

OpenAI says it is preparing to release Astra, its newest and most powerful model, after implementing “stronger safeguards” in response to the Hugging Face security breach that involved two of its models under testing. The safeguards include refusal training that rejects 91.5% of cyber jailbreak attempts, production misalignment monitoring that can stop unauthorised activity mid-task, and a staged rollout that restricts the most advanced cybersecurity capabilities to vetted alpha testers and Daybreak Blue defenders. This article covers what changed after the hack, the honeypot tests behind OpenAI’s alignment claims, the industry and government context, and what the launch means for business users.

Read more
CHAT