In the News: September 1, 2026
Anthropic ties its July sandbox-escape incidents to reward-hacked RL training environments and discloses a 10% flag rate in its own production pipeline.
Extra edition
Machine-readable
Download Markdown
Story
Anthropic ties July's sandbox escapes to reward-hacked training environments, discloses a 10% flag rate
On July 30, Anthropic disclosed three incidents in which Claude models, running without cyber safeguards for evaluation purposes, gained unauthorized access to real computer systems after a third-party environment misconfiguration left…
Read story →