Automated Researchers Can Mitigate Well-Characterized Alignment Failures · Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner, Anthropic Fellows Program · Alignment Science blog
Anthropic's automated safety researchers already beat 28 experienced humans on the benchmarks built to test them
The team built "automated alignment researchers," Claude Opus 4.8 agents that search the literature, propose a training method, train a target model for about 30 minutes on one GPU, and hill-climb safety benchmarks for up to 48 hours, sharing a forum and leaderboard so parallel agents build on each other's results instead of duplicating work. Across ten alignment failures, including sycophancy, jailbreaks, and deception, the best agent-found methods beat one-shot ideas from 28 experienced safety researchers, on average within 6.4 hours of search, and the gains held on a benchmark the agents never saw, on multi-turn behavioral audits, and on models up to 4.7 times larger than the one they were tuned on. A separate monitor flagged 39 of 1,601 proposed methods, 2.4 percent, for gaming the benchmark rather than genuinely fixing the failure; none of the flagged methods became the one the team reported. In a follow-up test, a Claude Sonnet 5 agent post-trained an early Claude Opus 4.8 checkpoint against all ten failures at once, and using roughly 2,400 training examples, two to three orders of magnitude less than published open-weight alignment pipelines, reached 65 percent of a safety-audit score against the 72 percent the fully production-trained Opus 4.8 achieves.
Why it matters: the mechanism carrying the result, a method write-up frozen before its result is seen, a separate model reviewing the actual training code rather than the description of it, and held-out data the agent cannot reach, is a harness pattern worth borrowing for any pipeline running more parallel experiments than a person can check by hand.