Overview
- Anthropic announced on Friday, August 7 that starting August 14 auto mode will be the default permission setting for Pro, Max, and Team users unless an administrator pins another setting.
- Auto mode uses a two‑stage classifier that runs a fast token filter then a deeper chain‑of‑thought check for flagged actions and enforces hard‑deny rules for data exfiltration and secret misuse.
- In Anthropic’s internal tests with 1,053 paid participants the company reported humans caught 13.6% of dangerous commands versus 89% for auto mode, and red‑teaming plus third‑party checks reduced missed attacks to single‑digit rates.
- Enterprise customers and the Claude API and cloud platforms remain opt‑in for now, administrators can override defaults, Anthropic will not bill the extra classifier tokens, and some teams report roughly 25% more pull requests when using auto mode.
- Anthropic warns the classifier lowers but does not eliminate risk and recommends human review for high‑stakes production changes while noting auto mode reduces approval fatigue and lets agents run longer multi‑step workflows.