Particle.news

OpenAI Flags Astra as First Model to Reach 'Critical' Cybersecurity Level

OpenAI restricted Astra after internal tests showed the model can discover and chain zero-day exploits, so the company will limit its strongest offensive features to a small set of vetted partners.

Overview

  • OpenAI announced on Tuesday that Astra meets its Preparedness Framework’s ‘Critical’ cyber threshold, meaning the model can autonomously find previously unknown vulnerabilities and develop end-to-end exploit chains.
  • In internal evaluations OpenAI says Astra scored 100 percent on ExploitBench, discovered two zero-day flaws while building a multi-stage exploit, and escaped a hardened browser sandbox in controlled tests.
  • OpenAI paused parts of Astra’s development for weeks, added isolated testing environments, increased agent monitoring, trained the model to refuse harmful cyber requests, and built a ‘misalignment monitor’ to stop unauthorized actions.
  • At launch OpenAI will keep Astra’s most advanced cyber capabilities limited to a small group of trusted Daybreak partners and select government and infrastructure organizations while broader defensive access is phased in.
  • Security teams warn the company’s safeguards may sometimes block legitimate vulnerability hunts or incident responses, and the Astra disclosure intensifies industry and policy debate about reporting standards, isolation, and identity-first controls after July’s Hugging Face testing breach.