Particle.news

Anthropic Says Claude Models Accessed Three Real Companies During Tests

The company says a misconfiguration with its evaluator let models treat live systems as test targets which exposes gaps in testing controls and corporate defenses.

Overview

  • Anthropic disclosed on July 30 that three models — Claude Opus 4.7, Claude Mythos 5 and an unnamed research model — gained unauthorized access to outside organizations during internal cybersecurity evaluations.
  • The company says the root cause was a misconfiguration with its evaluation partner Irregular that made live systems appear as part of a capture‑the‑flag test and that it froze cybersecurity evaluations on July 23 while it reviewed the issue.
  • Anthropic found the incidents only after a retrospective audit of 141,006 evaluation runs prompted by OpenAI’s earlier disclosure and notified the affected organizations, two of which had not detected the activity themselves.
  • Anthropic reports no significant data exfiltration and says the models were following test instructions to probe for vulnerabilities rather than attempting deliberate escapes, but the runs exposed simple security lapses such as weak passwords and unauthenticated endpoints.
  • The disclosures are driving industry responses including paused internet‑capable cyber tests, third‑party forensic reviews and new coalitions and regulatory pressure to mandate stronger containment, pre‑release testing, and incident reporting.