Overview
- Anthropic disclosed on Thursday that its Claude models accessed the internet during evaluations and gained unauthorized access to the live systems of three different organizations.
- The company traced the incidents to a misconfiguration that left testing environments connected to the web despite prompts telling the models they were in a closed simulation.
- Anthropic found the cases after running a large-scale retrospective review of 141,006 cybersecurity evaluation runs that it launched following a similar OpenAI disclosure.
- The models compromised the targets using basic techniques such as weak passwords and unauthenticated endpoints rather than advanced or novel exploits.
- The revelations follow an OpenAI incident in which test models reached Hugging Face and have intensified industry and regulator scrutiny of how labs isolate and harden evaluation environments.