The authors hand-coded 455 policy provisions from 102 communities into four rule types (Refuse, Disclose, Verify, Handoff) and built RepoComplianceBench: 106 issue instances from 49 repositories, run against four frontier agent and model…
In the News: August 2, 2026
A benchmark of 280 runs per agent finds coding agents obey rules that add a step and never obey rules that ask them to stop.
Coding agents obey rules that add work and ignore rules that ask them to stop
DeepSeek's price and rate-limit tables, read first-hand, and sixteen harness integrations
First-party figures per million tokens, as published: deepseek-v4-flash at $0.14 input on a cache miss, $0.0028 on a cache hit, $0.28 output, with a 2,500 concurrency limit; deepseek-v4-pro at $0.435, $0.003625 and $0.87, with a 500…
"I cannot be Linus and an LLM cannot be my subsystem maintainer"
A short note, labelled as such by the author, rebutting Antirez's framing of the expert developer as Linus with agents as his subsystem maintainers. Maganti's objection is trust, and specifically which decisions trust has to cover: "just…
The counterparty: stay in the loop and be outcompeted
Borretti extends the lights-out thesis from production to comprehension and frames it as a competitive fact rather than a preference: "No human needs to understand any of it, firms that put humans in the loop to understand the results…
On the Invalidity of the Claimed Disproof of Connes' Rigidity Conjecture
Only the abstract of this 19-page preprint is read here. Nielsen asserts a recent claimed disproof is invalid, having traced it "through 37,000 lines of published Lean code," and locates the failure in his own words: the disputed step "is a specification question the Lean kernel does not adjudicate." That is items 1 and 3's boundary, in the setting where verification was supposed to be total. A named claim with a machine-checked artifact behind it, not a settled result; no response from the original authors was looked for. Its Hacker News thread was flagged, at 31 points and 42 comments at 9.5 hours (read 09:12 EDT).
Codeberg Divides
Reports Codeberg has changed its terms to exclude projects "largely written with generative AI," and calls the decision legitimate but the wording unenforceable: "what does 'mostly' mean, and who can still tell?" Codeberg's terms text is unread here, so the change is reported as Ronacher describes it. Pair it with item 1: a platform-scale Refuse rule against 0% measured compliance.
We ran Kimi K3 through 3 more agent harnesses
Six harnesses, 26 to 28 identical tasks, model held constant. "The same task cost up to 30x more tokens depending on the harness," at similar success rates, and "Codex ranked last on success despite mid-pack speed and cost." Vendor-produced by a harness-tooling company, unreplicated, run counts and harness versions undisclosed.