---
title: 'In the News: October 6, 2026 (Extra 5)'
description: 'Two author-run preprints report that harness machinery moves coding-agent pass rates less than expected, while the bill moves a lot.'
canonical_url: 'https://darkfactory.dev/news/2026-10-06-extra-5'
markdown_url: 'https://darkfactory.dev/news/2026-10-06-extra-5.md'
collection: news
date_published: '2026-10-06T13:40:00-04:00'
date_modified: '2026-10-06T13:40:00-04:00'
---

# In the News: October 6, 2026 (Extra 5)


Two preprints from the past week report that harness machinery buys less pass rate than its cost suggests: one on SWE-bench Verified, one on ML-engineering tasks. Both are author-run, and both report the cost side.

## 1. Swapping coding-agent harnesses changed the bill more than the pass rate

**[What Does a Harness Buy? Tokens, Mostly](https://arxiv.org/abs/2610.04433)** · Yangze Liu, Zhongyi Han, Shandong University · arXiv preprint (v1), October 3, 2026

The authors held the model fixed and ran three harnesses as shipped, Claude Code, mini-SWE-agent and OpenCode, on 492 of the 500 SWE-bench Verified tasks, with a 300-step budget and no network access. On a pool of 447 tasks with two Qwen models and one run per cell, Claude Code and mini-SWE-agent were statistically equivalent within 5 points (differences of -1.8 and -1.4 points). On the 45 hardest tasks across five models, the spread between harnesses was 2 to 5 tasks per model, and rerunning one harness moved it by up to 3. Swapping the harness flipped 13% of tasks at the median, the same as rerunning the same harness; swapping the model flipped 22%. The one effect that cleared the noise was a loss: OpenCode trailed by up to about 9 points on the pool, and the authors trace about half of the gap on one model to runs cut off by an output cap with no recovery prompt. Cost per task differed by up to 3 times on the same model, which they tie to the fixed per-call preamble: 16,581 tokens for Claude Code, 7,025 for OpenCode and 829 for mini-SWE-agent. In their words, "Every harness effect we can name is a way to lose a task, and none we measured is a way to win one."

The results come from one benchmark and mostly open-weight or vendor-API models. Claude Opus 5 ran only inside Claude Code (36 of 45 hard tasks over three runs), so it gives no cross-harness comparison. The authors disclose unequal wall-clock limits that cost OpenCode six trials on one model's hard set, and they ran all baselines themselves. We read the main text and the first appendix; the paper says trial tables and analysis scripts are released.

**Why it matters:** The authors' own control is cheap to copy: rerun the same configuration before crediting a harness for a gap. By their power calculation, 45 tasks catch a gap of about 13 points only half the time, so small-sample harness comparisons deserve suspicion.

## 2. A single long agent session matched or beat four ML-engineering harnesses

**[How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?](https://arxiv.org/abs/2609.40303)** · Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel Abbé, authors at EPFL and Apple · arXiv preprint (v1), September 30, 2026

Built on OpenCode, the study varies one family of intervention at a time on MLE-bench and NatureBench tasks. The largest effect was giving the model a coding-agent environment, a shell and a filesystem, in place of a chat interface. Beyond that, the authors report no added machinery with a statistically significant gain. Their baseline, a single session with three small tools and a prompt to continue, earned a medal on 62.5% of competitions with GLM 5.2, against 47.1% for the best of four open-source harnesses they ran under the same time budget. Multi-agent additions did not help on their fixed 14-task set. The authors' modelled cost for the single session was $12.12 against $1.95 for one competing harness, about 6.2 times as much.

This is ML engineering, not software engineering, and the authors built the baseline and ran the competing harnesses themselves, acknowledging possible asymmetry in tuning effort. The seed count per cell was not stated in the main text we read, and we did not read the appendices. The paper concludes that "The leverage is in the model and in the runtime it is given, not in the scaffolding built around them."

**Why it matters:** It points the same way as the item above from a different domain: test a single long session with a real execution environment as the baseline before building orchestration. The cost figure is the caveat, since that baseline was roughly six times as expensive as one alternative.
