Fighting Code Slop: The State of Software Factories · Dex Horthy, CEO and co-founder, HumanLayer · Conference talk, Agentic AI Foundation
Horthy argues unattended models will not keep a codebase healthy
Horthy, who says his company sells a multiplayer coding agent workspace, opens by arguing that unattended models will not improve or maintain codebase quality over time, so teams should keep reading the code for now. He cites a report he places in May saying that since January code review quality has fallen and incidents per pull request and bugs per developer have risen; the talk does not name it. He also cites SlopCodeBench, from a University of Wisconsin lab, which reveals a task in stages so a model has to keep extending its own code. He says the best model at launch scored 14.8% (around 7:00). His explanation is that reinforcement learning needs a fast check, and maintainability has none, because bad architecture shows up months later.
His practices include planning before building without over-planning, giving agents a browser and tools such as curl to test their work, and a roughly 100-rule set of anti-slop lint rules for TypeScript from Dylan Moloy. He also suggests attaching an agent-written fix to the page when something breaks at night, and says that merging even half of those would raise throughput. The captions behind this item are auto-generated and were not checked against the audio, so the figures and names above come from a machine transcript.
Why it matters: This is a practitioner's argument with some borrowed data, and he sells a product in the same area. The two cited sources are his citations, not our checks. The practical content is the list of controls, and it lines up with the other two items: review gates, a check the agent cannot rewrite, and tests the agent runs itself.