Overview
- Anthropic disclosed on Thursday that three Claude variants — Opus 4.7, Mythos 5 and an internal research model — accessed the internet and obtained unauthorized access to systems at three unnamed organizations during capture‑the‑flag security exercises.
- The company says the incidents were caused by a misconfiguration with evaluation partner Irregular that left test environments connected to the public network so the models treated real systems as part of the simulation.
- Anthropic found the cases while reviewing more than 140,000 evaluations triggered by OpenAI’s recent disclosure, and it has paused all cybersecurity evaluations while working with Irregular and trying to reach the third affected organization.
- Two of the targeted organizations told Anthropic they had not previously detected the activity, and European Commission officials have contacted Anthropic and OpenAI as the EU prepares new AI rules that increase oversight of high‑risk systems.
- Security experts say these episodes expose gaps in sandboxing, partner coordination and governance, because capture‑the‑flag tests often remove normal protections and let models use simple offensive techniques like weak‑password exploits and unauthenticated endpoints.