---
title: 'In the News: October 11, 2026 (Extra 2)'
description: 'Cockroach Labs reports five months of agents behind human approval, and Dex Horthy argues unattended models will not keep a codebase healthy.'
canonical_url: 'https://darkfactory.dev/news/2026-10-11-extra-2'
markdown_url: 'https://darkfactory.dev/news/2026-10-11-extra-2.md'
collection: news
date_published: '2026-10-11T11:33:13-04:00'
date_modified: '2026-10-11T11:33:13-04:00'
---

# In the News: October 11, 2026 (Extra 2)


Cockroach Labs says it ran coding agents behind a plan-and-review process for five months and turned on human approval after the first week. Its account also puts a price on that process. Separately, a conference talk from HumanLayer's CEO argues that models do not keep a codebase healthy without supervision, and a decompilation project shows one way to give agents a check they cannot easily bend.

## 1. Cockroach Labs reports five months of agents working behind review gates

**[Five months treating bugs like patients and coding agents like a medical team](https://www.cockroachlabs.com/blog/experiment-running-hospital-code/)** · Adam Storm and Rafi Shamim, Cockroach Labs · Cockroach Labs Blog, October 9, 2026

The authors describe agents that take issues through a workup, a plan and a review before any code is written, from April 21 to September 11. Their summary rule is "No code before a plan, no plan without review." After the first week they turned on a human approval mode that requires a person to look at every merge. They say the pipeline handled "over a million lines of code with only 7 total reverts," that they can reconstruct why any of 1,238 pull requests was merged, and that the run cost just over $135,000 in Claude tokens, about $84 for the average patient, their term for an issue. The Migration Assistant, now in preview for Postgres, was in their words almost entirely written this way. In an earlier Db2 test, they write, "we weren't reviewing any of the code that the hospital wrote during that initial test," and they list confirming its correctness as ongoing work.

The process also got heavy. A July audit of 25 skill files, a little over 100,000 words, flagged about 23,000 words as removable without changing any gate, command or template. One fix that came to a single line took eleven rework rounds over two days. They now hand a pull request to an Attending reviewer after four consecutive rework rounds that each turned up a new blocking finding, or after six substantive rounds of any kind.

**Why it matters:** This is a company's account of its own repository, so the totals are its figures and nothing here tests them independently. What a team can borrow is concrete: a plan-review gate before code, human approval on every merge, and numeric thresholds that stop a review loop from running indefinitely. Whether those thresholds work as intended is not shown yet.

## 2. Heumann protects the checker after agents produce plausible, wrong code

**[500+ Billion Tokens Later: Letting AI Agents Decompile A First-Person Shooter](https://momo5502.com/posts/2026-10-09-game-decompilation/)** · Maurice Heumann · Maurice's Blog, October 9, 2026

Heumann says agents decompiled about 80% of a game in four weeks, enough to launch it, reach the main menu and load maps. The code had wrong function signatures, types and struct layouts, and his post says agents also invented or removed logic. His fix was a byte-level check: a reconstructed object file is compared with the original executable and debug database, and a function counts as exact only if they match. Relocations are excluded from the comparison and checked by symbol and offset instead. He also protects the check itself: CI hashes the verification script and compares it with a stored GitHub Actions secret.

After further work he reports that 99% of the game's functions are present in the reconstructed source and 83% are byte exact. The code stays private, so no one can reproduce those figures. The token count is an estimate of 600 to 700 billion, and he says lost session logs prevent an exact total. His lesson: "Correctness should be defined and machine-checkable."

**Why it matters:** The approach depends on having an original binary to compare against, which most software work does not. Where an external reference does exist, this shows the check needs to be outside the agent's reach. The 83% is a measure of matching, not a finding about the other 17%, and no one outside the project can inspect the code.

## 3. Horthy argues unattended models will not keep a codebase healthy

**[Fighting Code Slop: The State of Software Factories](https://www.youtube.com/watch?v=ix1qQK1IvmA)** · Dex Horthy, CEO and co-founder, HumanLayer · Conference talk, Agentic AI Foundation

Horthy, who says his company sells a multiplayer coding agent workspace, opens by arguing that unattended models will not improve or maintain codebase quality over time, so teams should keep reading the code for now. He cites a report he places in May saying that since January code review quality has fallen and incidents per pull request and bugs per developer have risen; the talk does not name it. He also cites SlopCodeBench, from a University of Wisconsin lab, which reveals a task in stages so a model has to keep extending its own code. He says the best model at launch scored 14.8% (around 7:00). His explanation is that reinforcement learning needs a fast check, and maintainability has none, because bad architecture shows up months later.

His practices include planning before building without over-planning, giving agents a browser and tools such as curl to test their work, and a roughly 100-rule set of anti-slop lint rules for TypeScript from Dylan Moloy. He also suggests attaching an agent-written fix to the page when something breaks at night, and says that merging even half of those would raise throughput. The captions behind this item are auto-generated and were not checked against the audio, so the figures and names above come from a machine transcript.

**Why it matters:** This is a practitioner's argument with some borrowed data, and he sells a product in the same area. The two cited sources are his citations, not our checks. The practical content is the list of controls, and it lines up with the other two items: review gates, a check the agent cannot rewrite, and tests the agent runs itself.
