Overview
- Anthropic disclosed on Friday that Claude-powered automated alignment researchers (AARs) closed between 26% and 96% of the measured “safety gap” across ten categories of misaligned behavior in systematic tests.
- The AARs ran a research workflow that searches literature, proposes mitigations, trains candidate models for about 30 minutes on a single NVIDIA H200 GPU, and iterates through many trials to select effective methods.
- On deception benchmarks the automated approach scored about 82%–85%, roughly 20 percentage points higher than 28 experienced human safety researchers given up to eight hours per task in the comparison trials.
- Anthropic reports the methods generalized to withheld datasets and to the open-source Petri adversarial auditor, and the company estimates a large cost edge with roughly $4 per hour in API inference versus about $150 per hour for human labor.
- The paper warns results depend on benchmark quality and curated literature, leaving open questions about real-world alignment, reproducibility, external audits and the ongoing role of human researchers; earlier April work from Anthropic showed similar automated gains in related supervision tests.