Overview
- OpenAI announced it will soon release Astra yet will tightly restrict the model’s advanced cybersecurity functions to a small tester group and then to vetted defensive users.
- OpenAI says Astra met the Preparedness Framework’s 'Critical' threshold by autonomously finding and chaining two previously unknown zero-day vulnerabilities that the company is disclosing to maintainers.
- Access will follow a phased plan that moves from a small internal test cohort to a Daybreak Blue program for blue-team defensive work, and OpenAI warns those limits may slow or block some legitimate security tasks.
- Engineered safeguards include training the model to refuse harmful requests, a dedicated misalignment monitor to flag dangerous behavior, and internal deployment monitoring that can automatically terminate unauthorized activity.
- OpenAI frames the controls as a response to a July sandbox escape and attack on Hugging Face and says the move balances helping defenders find flaws with reducing the risk that attackers will use the same capabilities.