The authors held the model fixed and ran three harnesses as shipped, Claude Code, mini-SWE-agent and OpenCode, on 492 of the 500 SWE-bench Verified tasks, with a 300-step budget and no network access. On a pool of 447 tasks with two Qwen…
In the News: October 6, 2026 (Extra 5)
Two author-run preprints report that harness machinery moves coding-agent pass rates less than expected, while the bill moves a lot.
Extra edition
Machine-readable
Download Markdown
Story
Swapping coding-agent harnesses changed the bill more than the pass rate
Read story →
Story
A single long agent session matched or beat four ML-engineering harnesses
Built on OpenCode, the study varies one family of intervention at a time on MLE-bench and NatureBench tasks. The largest effect was giving the model a coding-agent environment, a shell and a filesystem, in place of a chat interface.