← In the News

Swapping coding-agent harnesses changed the bill more than the pass rate

What Does a Harness Buy? Tokens, Mostly · Yangze Liu, Zhongyi Han, Shandong University · arXiv preprint (v1), October 3, 2026

Machine-readable Download Markdown

The authors held the model fixed and ran three harnesses as shipped, Claude Code, mini-SWE-agent and OpenCode, on 492 of the 500 SWE-bench Verified tasks, with a 300-step budget and no network access. On a pool of 447 tasks with two Qwen models and one run per cell, Claude Code and mini-SWE-agent were statistically equivalent within 5 points (differences of -1.8 and -1.4 points). On the 45 hardest tasks across five models, the spread between harnesses was 2 to 5 tasks per model, and rerunning one harness moved it by up to 3. Swapping the harness flipped 13% of tasks at the median, the same as rerunning the same harness; swapping the model flipped 22%. The one effect that cleared the noise was a loss: OpenCode trailed by up to about 9 points on the pool, and the authors trace about half of the gap on one model to runs cut off by an output cap with no recovery prompt. Cost per task differed by up to 3 times on the same model, which they tie to the fixed per-call preamble: 16,581 tokens for Claude Code, 7,025 for OpenCode and 829 for mini-SWE-agent. In their words, "Every harness effect we can name is a way to lose a task, and none we measured is a way to win one."

The results come from one benchmark and mostly open-weight or vendor-API models. Claude Opus 5 ran only inside Claude Code (36 of 45 hard tasks over three runs), so it gives no cross-harness comparison. The authors disclose unequal wall-clock limits that cost OpenCode six trials on one model's hard set, and they ran all baselines themselves. We read the main text and the first appendix; the paper says trial tables and analysis scripts are released.

Why it matters: The authors' own control is cheap to copy: rerun the same configuration before crediting a harness for a gap. By their power calculation, 45 tasks catch a gap of about 13 points only half the time, so small-sample harness comparisons deserve suspicion.