Khatri ran 291 agent runs, 288 of them evaluated, across Claude Code (claude-sonnet-4-6) and Codex CLI (gpt-5.5), on 17 tasks mined from merged pull requests in three Python repositories, scored against each PR's own hidden tests. Three…
In the News: August 4, 2026
A controlled two-agent ablation finds AGENTS.md and CLAUDE.md do not measurably move coding-agent correctness on either Claude Code or Codex.
A controlled ablation finds AGENTS.md does not move correctness on either frontier agent
Models systematically avoid deleting code, and the tests that pass them do not check
This item is written from the paper's abstract. The body is unread by this edition.
Claude Code 2.1.221 closes two permission-check bypasses and adds credential masking
Today's release fixes, in Anthropic's own words, "a Bash tool permission-check bypass where zsh could execute hidden commands in [[ ]] regex conditionals; affected commands now prompt for permission," plus a PowerShell permission check that…
Goedecke: the human is the bottleneck, not the model
Goedecke's claim is that the most important prompting skill is domain expertise in the thing being prompted about, and his evidence is a specific artifact: a public ChatGPT transcript of Terence Tao working on a counterexample to the…
Elon Musk: "Source code is on the verge of becoming like assembly"
Musk's post, dated 3 August, argues the next step is "getting rid of 'source code' entirely and just making an efficient binary directly with AI." Dex Horthy of HumanLayer, whose Why Software Factories Fail is the most-cited artifact in this lane, quote-tweeted it to disagree: "It is the future but we are far from being 'on the verge' of it for serious software with uptime." Horthy's post stood at 270 favourites and 33K views at roughly 18 hours. Its permalink was not captured, so it is reached here through Musk's post.
Gordon Mickel claims a software factory that does not fail
Mickel, who leads AI at GrowthFactors, says he has been running his own factory setup against SlopCodeBench and "they don't need to fail." A write-up has been promised twice and is not published. Reported as a claim by a named practitioner, not a result: no methodology, no numbers, no artifact yet. At 8 favourites and 8.4K views at 14 hours it is travelling far less well than Horthy's reply to it.