Overview
- OpenAI said Friday it paused internal Astra work that did not meet strengthened controls after preliminary tests showed the model may reach its highest 'Critical' cybersecurity tier.
- Under OpenAI’s Preparedness Framework, 'Critical' means a model could autonomously find and build working zero‑day exploits for hardened systems or plan and execute novel end‑to‑end cyberattacks from a high‑level goal.
- As an immediate step, OpenAI moved Astra into isolated environments with restricted network and tool access, encrypted model weights, sandboxed execution, and real‑time monitoring that can interrupt risky agent actions.
- OpenAI confirmed Astra was not involved in the earlier Hugging Face breach and said it will give selected regulators and safety organisations access to run independent tests before any broader release.
- The disclosure follows multiple recent containment failures at frontier labs and has accelerated industry efforts to build shared defensive tools, formal testing protocols, and tighter rules for third‑party evaluators.