← In the News

A single long agent session matched or beat four ML-engineering harnesses

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering? · Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel Abbé, authors at EPFL and Apple · arXiv preprint (v1), September 30, 2026

Machine-readable Download Markdown

Built on OpenCode, the study varies one family of intervention at a time on MLE-bench and NatureBench tasks. The largest effect was giving the model a coding-agent environment, a shell and a filesystem, in place of a chat interface. Beyond that, the authors report no added machinery with a statistically significant gain. Their baseline, a single session with three small tools and a prompt to continue, earned a medal on 62.5% of competitions with GLM 5.2, against 47.1% for the best of four open-source harnesses they ran under the same time budget. Multi-agent additions did not help on their fixed 14-task set. The authors' modelled cost for the single session was $12.12 against $1.95 for one competing harness, about 6.2 times as much.

This is ML engineering, not software engineering, and the authors built the baseline and ran the competing harnesses themselves, acknowledging possible asymmetry in tuning effort. The seed count per cell was not stated in the main text we read, and we did not read the appendices. The paper concludes that "The leverage is in the model and in the runtime it is given, not in the scaffolding built around them."

Why it matters: It points the same way as the item above from a different domain: test a single long session with a real execution environment as the baseline before building orchestration. The cost figure is the caveat, since that baseline was roughly six times as expensive as one alternative.