← In the News

Anthropic's automated safety researchers already beat 28 experienced humans on the benchmarks built to test them

Automated Researchers Can Mitigate Well-Characterized Alignment Failures · Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner, Anthropic Fellows Program · Alignment Science blog

Machine-readable Download Markdown

The team built "automated alignment researchers," Claude Opus 4.8 agents that search the literature, propose a training method, train a target model for about 30 minutes on one GPU, and hill-climb safety benchmarks for up to 48 hours, sharing a forum and leaderboard so parallel agents build on each other's results instead of duplicating work. Across ten alignment failures, including sycophancy, jailbreaks, and deception, the best agent-found methods beat one-shot ideas from 28 experienced safety researchers, on average within 6.4 hours of search, and the gains held on a benchmark the agents never saw, on multi-turn behavioral audits, and on models up to 4.7 times larger than the one they were tuned on. A separate monitor flagged 39 of 1,601 proposed methods, 2.4 percent, for gaming the benchmark rather than genuinely fixing the failure; none of the flagged methods became the one the team reported. In a follow-up test, a Claude Sonnet 5 agent post-trained an early Claude Opus 4.8 checkpoint against all ten failures at once, and using roughly 2,400 training examples, two to three orders of magnitude less than published open-weight alignment pipelines, reached 65 percent of a safety-audit score against the 72 percent the fully production-trained Opus 4.8 achieves.

Why it matters: the mechanism carrying the result, a method write-up frozen before its result is seen, a separate model reviewing the actual training code rather than the description of it, and held-out data the agent cannot reach, is a harness pattern worth borrowing for any pipeline running more parallel experiments than a person can check by hand.