arXiv, Aug 18, 2026. Rewriting a codebase in ways that preserve its behavior (renaming variables, adding dead code) dropped agent resolve rates by up to 6.7 percentage points in the worst configurations, with no single model's robustness ranking holding across two different agent scaffolds. Which harness wraps the model may matter as much as which model it is.
In the News: August 27, 2026
A five-person team banned human-written pull requests and published the numbers behind it: 295 issues, 217 shipped, in 30 days.
Morning edition
Machine-readable
Download Markdown
Story
A five-person company banned human-written pull requests and published the numbers
In January, Paul Stack's team threw away a Rust codebase they had spent years building and rebuilt their process around agents from the start. The new rule: agents write every line of code, and a human-written pull request does not get…
Read story →
Story
Why an agent's answer can look right and still be wrong
A healthcare analytics agent was asked for Medicare Advantage readmission rates tied to a specific diagnosis and date range. It returned a plausible number that was wrong three separate ways: it mixed in the wrong plan type, used an…
Read story →
Also this cycle
Permalink
Mahmud, Gupta, Chaudhary, Enis, Mangal, Singh and Pasareanu (Colorado State, Microsoft, UIUC, Carnegie Mellon)