Particle.news

Anthropic’s Claude Automates Fixes for Alignment Failures

Anthropic says its automated researchers can propose, test and deploy short training fixes that cut measured safety gaps at far lower cost than human teams.

Overview

  • Anthropic disclosed on Friday that Claude-powered automated alignment researchers (AARs) closed between 26% and 96% of the measured “safety gap” across ten categories of misaligned behavior in systematic tests.
  • The AARs ran a research workflow that searches literature, proposes mitigations, trains candidate models for about 30 minutes on a single NVIDIA H200 GPU, and iterates through many trials to select effective methods.
  • On deception benchmarks the automated approach scored about 82%–85%, roughly 20 percentage points higher than 28 experienced human safety researchers given up to eight hours per task in the comparison trials.
  • Anthropic reports the methods generalized to withheld datasets and to the open-source Petri adversarial auditor, and the company estimates a large cost edge with roughly $4 per hour in API inference versus about $150 per hour for human labor.
  • The paper warns results depend on benchmark quality and curated literature, leaving open questions about real-world alignment, reproducibility, external audits and the ongoing role of human researchers; earlier April work from Anthropic showed similar automated gains in related supervision tests.