---
title: 'In the News: September 26, 2026 (Morning)'
description: "A study of Claude's own coding-session claims finds roughly one in four to five responses contain a wrong claim, worst on proposed fixes."
canonical_url: 'https://darkfactory.dev/news/2026-09-26-morning'
markdown_url: 'https://darkfactory.dev/news/2026-09-26-morning.md'
collection: news
date_published: '2026-09-26T07:20:00-04:00'
date_modified: '2026-09-26T07:20:00-04:00'
---

# In the News: September 26, 2026 (Morning)


A single new paper carries this edition: the first measured error rate for what a coding agent says to its user, separate from the code it writes.

## 1. A study of Claude's own claims to its user finds it wrong about one time in four

**[Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase](https://arxiv.org/abs/2609.29744)** · Douglas Leith, Trinity College Dublin · arXiv, Sep 24, 2026

Leith built a fact-checking pipeline and ran it against 678 sessions of Claude building a real 21,000-line tool with no human-written code anywhere in its history. Checking a stated fact held up best, at 94.3% accurate. Explaining how existing code works came in at 91.1%, diagnosing a bug's root cause at 89.3%, and proposing a fix for a bug did worst, at 79.4%. Across a full response, roughly one in four to five contained at least one wrong claim. Mining the session transcripts, rather than relying on commit history alone, also found that 14.3% of code-generation events introduced a real bug that the AI's own test suite caught before anything was committed, a rate no prior commit-based study could see. In one quoted example, an agent diagnosed a missing file as living "in my tool sandbox, not your shell's filesystem," then ran the identical command itself from that same sandbox, hit the identical error, and moved on to a workaround without revising the diagnosis.

**Why it matters:** The paper's underlying dataset, scoring rubric, and regeneration scripts are public, so these figures can be rerun rather than taken on faith. A proposed fix is the category most worth slowing down to verify, since it is the one most likely to be wrong.
