Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase · Douglas Leith, Trinity College Dublin · arXiv, Sep 24, 2026
A study of Claude's own claims to its user finds it wrong about one time in four
Leith built a fact-checking pipeline and ran it against 678 sessions of Claude building a real 21,000-line tool with no human-written code anywhere in its history. Checking a stated fact held up best, at 94.3% accurate. Explaining how existing code works came in at 91.1%, diagnosing a bug's root cause at 89.3%, and proposing a fix for a bug did worst, at 79.4%. Across a full response, roughly one in four to five contained at least one wrong claim. Mining the session transcripts, rather than relying on commit history alone, also found that 14.3% of code-generation events introduced a real bug that the AI's own test suite caught before anything was committed, a rate no prior commit-based study could see. In one quoted example, an agent diagnosed a missing file as living "in my tool sandbox, not your shell's filesystem," then ran the identical command itself from that same sandbox, hit the identical error, and moved on to a workaround without revising the diagnosis.
Why it matters: The paper's underlying dataset, scoring rubric, and regeneration scripts are public, so these figures can be rerun rather than taken on faith. A proposed fix is the category most worth slowing down to verify, since it is the one most likely to be wrong.