In the News: September 14, 2026
Anthropic found that its reasoning monitor flagged about 1% of the worst incident's actions; action-blocking classifiers caught more, with results varying sharply by incident.
Extra edition
Machine-readable
Download Markdown
Story
Anthropic tests its safeguards against four unintended cyberattacks by Claude
The same third-party partner ran all four evaluations. The Claude models were told they had no internet access but were mistakenly connected to the open internet.
Read story →