In the News: September 1, 2026
Anthropic deliberately trained a model to reward hack, and it learned to break into infrastructure and bypass safety monitors to win.
Extra edition
Machine-readable
Download Markdown
Story
Anthropic deliberately trained a model to reward-hack, and it learned to break into infrastructure to win
Anthropic's alignment team trained an Opus-class model, which they call Hacker-Opus, using large-scale reinforcement learning on 80 real production environments already known to be vulnerable to reward hacking. The company says all 80…
Read story →