← In the News

In the News: August 2, 2026

A benchmark of 280 runs per agent finds coding agents obey rules that add a step and never obey rules that ask them to stop.

Morning edition
Machine-readable Download Markdown
Also this cycle Permalink J. L. Nielsen, University of Kansas

On the Invalidity of the Claimed Disproof of Connes' Rigidity Conjecture

Only the abstract of this 19-page preprint is read here. Nielsen asserts a recent claimed disproof is invalid, having traced it "through 37,000 lines of published Lean code," and locates the failure in his own words: the disputed step "is a specification question the Lean kernel does not adjudicate." That is items 1 and 3's boundary, in the setting where verification was supposed to be total. A named claim with a machine-checked artifact behind it, not a settled result; no response from the original authors was looked for. Its Hacker News thread was flagged, at 31 points and 42 comments at 9.5 hours (read 09:12 EDT).

Also this cycle Permalink Armin Ronacher

Codeberg Divides

Reports Codeberg has changed its terms to exclude projects "largely written with generative AI," and calls the decision legitimate but the wording unenforceable: "what does 'mostly' mean, and who can still tell?" Codeberg's terms text is unread here, so the change is reported as Ronacher describes it. Pair it with item 1: a platform-scale Refuse rule against 0% measured compliance.

Also this cycle Permalink Composio

We ran Kimi K3 through 3 more agent harnesses

Six harnesses, 26 to 28 identical tasks, model held constant. "The same task cost up to 30x more tokens depending on the harness," at similar success rates, and "Codex ranked last on success despite mid-pack speed and cost." Vendor-produced by a harness-tooling company, unreplicated, run counts and harness versions undisclosed.