arXiv, August 20, 2026 · A literature review of contradictory productivity claims about AI coding tools. It proposes a codebase's age as a testable variable behind the scatter and identifies a gap between audit-tier findings and the larger numbers that circulate publicly.
In the News: August 26, 2026
An independent Terminal-Bench 2.1 run shows a harness upgrade can match a newer model at a fraction of the cost, with exact transfer failing across providers.
Morning edition
Machine-readable
Download Markdown
Story
A harness upgrade matched a model-generation jump, with costs and failures on the record
The authors built StateM, a runtime that wraps a coding agent in durable states, checked transitions, and versioned runbooks without touching model weights. Applied to GPT-5.5 on Terminal-Bench 2.1, it raised accuracy from an 83.1 percent…
Read story →
Also this cycle
Permalink
Michels, Abu Ghazaleh, Lazzari, Kassem, and Klein, KAUST and collaborators