Overview
- Starting August 14, Anthropic will switch Auto mode on by default for Pro, Max and Team accounts so the agent proceeds without per‑step human approval unless an action is judged irreversible, destructive, or outside the user’s environment.
- Every tool call in Auto mode is routed through a safety classifier that blocks risky actions, tries safer alternatives, or asks for permission and will fall back to manual approvals after repeated blocks.
- Anthropic says controlled testing and independent evaluations found the classifier flagged far more dangerous commands than human reviewers — reporting Auto mode blocked 89% of harmful actions versus 13.6% for manual review in a 1,053‑tester study, and Trajectory Labs found zero success in 720 prompt‑injection attempts against Auto mode.
- The change is optional for enterprise, API and cloud deployments for now so administrators can evaluate it, and teams can pin defaults or disable Auto mode during the staged rollout.
- Anthropic removed classifier token charges for the affected paid tiers, cites customers who report higher developer throughput on Auto mode, and warns the feature reduces but does not eliminate risk so it recommends human review for high‑stakes production changes.