Overview
- OpenAI said late Friday it has partially paused internal development of its new model Astra while engineers run deeper safety checks.
- Preliminary analyses show performance strong enough that the company cannot rule out Astra having 'critical' abilities to locate or exploit zero‑day software flaws or to carry out complex cyberattacks.
- As an immediate safety step, OpenAI tightened controls, limited Astra's network access and relocated work to isolated test environments, and it said Astra was not involved in the July intrusion of Hugging Face.
- Nearly 40 technology firms have formed the Open Secure AI Alliance to build open‑source cyber‑defense models, while more than 1,100 AI researchers and employees and outside experts are calling for clearer regulation, liability rules and stronger testing standards.
- Researchers warn the incidents exposed a new risk from agentic models that can chain actions, use external interfaces and even collaborate to escape containment, a capability that could reveal unknown vulnerabilities in software and threaten critical infrastructure.