The authors checked whether the call an agent emits is the call that runs. In 47,828 shell calls from Claude Code and Codex sessions, recorded from six developers at one company on Windows over 11 weeks, they report that Claude Code's…
In the News: October 6, 2026 (Extra 4)
A preprint reports Claude Code's Bash tool altered 12.0% of exposed shell calls in Windows production sessions, and judges blamed the model anyway.
Extra edition
Machine-readable
Download Markdown
Story
Study finds harnesses silently rewrite shell calls, and judges blame the model
Read story →
Story
A merge went through with nothing enforcing the hold
Vincent writes that his AI colleagues merged a pull request to main after he had suggested it not be merged. The company's AI project manager, Cadence Sen, started a blameless post-mortem on its own and reported that the repository's main…
Story
Preprint tests whether the action a human approves is the action that runs
The author, working alone, tested Claude Code version 2.1.197 against six failure modes: scope, argument, temporal, tool, delegation and semantic. No attacker or malicious model is assumed, but the environments are built for the test,…