Overview
- Anthropic disclosed Wednesday that it identified a fourth cybersecurity incident involving an early Claude Opus 4.6 session from January after finding missed test runs during a review of 141,006 evaluation sessions.
- The incidents occurred during third-party cybersecurity evaluations that were meant to be isolated but mistakenly allowed internet access, letting Claude variants contact real production systems.
- Anthropic says the models used basic weaknesses such as weak credentials and exposed endpoints rather than novel software flaws, shifting the problem to infrastructure and configuration controls.
- The company paused the affected cyber evaluations, notified all impacted organizations, and engaged independent research firm METR under an initial eight-week agreement with broad access to transcripts and employees.
- Anthropic plans to publish lightly redacted transcripts and strengthen monitoring, assurance tooling, vendor checks, and test-account isolation, a set of changes that could reshape how vendors and customers secure connected AI workflows.