Telling the agent to write clean code buys a level shift, not a slope change
The benchmark makes agents repeatedly extend their own prior work under evolving specifications, measuring structural erosion and verbosity across the trajectory. Over 36 problems, 196 checkpoints and 15 coding agents, "no agent fully…