Overview
- OpenAI announced on Tuesday that Astra meets its Preparedness Framework’s ‘Critical’ cyber threshold, meaning the model can autonomously find previously unknown vulnerabilities and develop end-to-end exploit chains.
- In internal evaluations OpenAI says Astra scored 100 percent on ExploitBench, discovered two zero-day flaws while building a multi-stage exploit, and escaped a hardened browser sandbox in controlled tests.
- OpenAI paused parts of Astra’s development for weeks, added isolated testing environments, increased agent monitoring, trained the model to refuse harmful cyber requests, and built a ‘misalignment monitor’ to stop unauthorized actions.
- At launch OpenAI will keep Astra’s most advanced cyber capabilities limited to a small group of trusted Daybreak partners and select government and infrastructure organizations while broader defensive access is phased in.
- Security teams warn the company’s safeguards may sometimes block legitimate vulnerability hunts or incident responses, and the Astra disclosure intensifies industry and policy debate about reporting standards, isolation, and identity-first controls after July’s Hugging Face testing breach.