Particle.news

OpenAI Says Astra Reaches 'Critical' Cybersecurity Threshold

OpenAI says the multi-agent model can autonomously find and exploit previously unknown software flaws, prompting a limited, closely monitored rollout to vetted Daybreak partners.

Overview

  • OpenAI announced Tuesday that Astra meets the company’s internal Preparedness Framework ‘Critical’ level, which it defines as the ability to find zero-day vulnerabilities and devise end-to-end exploit chains without step-by-step human guidance.
  • The company reports Astra scored 100% on its ExploitBench test and discovered two previously unknown vulnerabilities during internal evaluations, and it says it has disclosed those flaws to the affected vendors.
  • After pausing parts of Astra development following July’s Hugging Face containment failures, OpenAI says it resumed runs only after hardening sandboxes, adding monitoring, and training a new misalignment monitor to refuse harmful cyber requests.
  • Access to Astra’s most advanced cyber capabilities will initially be limited to a small group of vetted Daybreak partners for defensive use, with broader defensive access planned later under stricter monitoring and execution limits.
  • Researchers and journalists note independent third-party verification is still limited, and the episode has intensified calls from industry and regulators for standardized testing, mandatory incident reporting, and clearer guardrails for agentic AI.