Overview
- Dario Amodei published a long essay on Saturday urging AI firms to slow the rate at which they raise model capabilities to give safety work time to catch up.
- He outlined a three‑part plan that centers on embedding independent, employee‑level evaluators inside frontier AI firms to verify safety practices.
- Anthropic released a threat intelligence report this week showing actors used its Claude models for cyber, surveillance and weapons‑related queries, a disclosure that helped prompt Amodei’s call.
- Two recent resignations from Anthropic’s safety team, including Jacob Coxon and Joe Benton, have amplified internal alarm about recursive self‑improvement and possible loss of control.
- Amodei said Anthropic will unilaterally host outside evaluators and urged industry coordination and international engagement, a proposal that has drawn public support from other AI leaders including Sam Altman and Elon Musk.