← In the News

Preprint gives every repository file its own model, with the biggest reported gains on whole-repository migration

Harness Engineering for Software Engineering via Modular Executable Dev-Primitives · Haibo Jin, Xinjie Li, Peng Kuang, Haohan Wang, University of Illinois Urbana-Champaign and Penn State · arXiv preprint (v1), October 6, 2026

Machine-readable Download Markdown

The authors wrap each repository component, such as a source, config or test file, in what they call a Dev-Primitive: the artifact paired with a resident LLM that reads and edits it and sends natural-language messages to the callers, callees, configs and tests it affects. Their framework, HERMES, activates only the primitives a task appears to need and runs the repository. If the tests still fail, a diagnosis step maps the failure back to the components it implicates and reactivates only those. The default revision budget is three rounds.

With GPT-5.6 Sol at medium reasoning effort, the authors report an average gain of 12.4 percentage points over matched baseline harnesses across four benchmarks. The gains are uneven. On SWE-bench Verified, GPT-5.6 Sol moves from 96.2% to 97.0% against mini-SWE-agent. On SWE Refactor Bench, a whole-repository migration benchmark, its composite score rises from 6.5% to 31.0% against a Codex baseline, and on Terminal-Bench 4.0 from 33.0% to 51.8%. Claude Opus 5 gains 1.5 and 1.2 points on Terminal-Bench 4.0 against Claude Code at the two effort levels tested. The paper says the migration gains "generally require higher inference cost": at medium effort GPT-5.6 Sol's reported per-task cost on SWE Refactor Bench goes from $6.0 to $24.8.

The cheaper-model result is the other headline. With Qwen3-8B as the Dev-Primitives and stronger models handling activation and diagnosis, the authors report staying within 4.5 percentage points of an all-GPT-5.6 Sol setup on all four benchmarks, at 26.2% lower inference cost on Terminal-Bench 4.0 ($2.51k against $3.40k). Their ablation puts more weight on diagnosis than on activation: using GPT-5.5 for diagnosis alone lifts Terminal-Bench 4.0 from 27.6% to 46.1%, against 36.7% for activation alone. The paper also states that replacing Dev-Primitives with a centralized editor over the same selected components was the most damaging change tested, costing 14.2 points on SWE Refactor Bench and 13.0 on Terminal-Bench 4.0.

These are the authors' own figures from a first-version preprint, and HERMES was run once per backbone. The authors report that three repeats of the default GPT-5.6 Sol setting on SWE Refactor Bench gave 31.0 ± 1.3. We read the main text, not the appendices, which hold the baseline provenance and the limitations section.

Why it matters: For anyone building a harness for multi-file work, the useful part is the ablation: the authors attribute the centralized-editor losses to failures that appear at the repository level with no component owning them to route back to. The cost column is the counterweight, since the migration gains come with roughly four times the per-task spend at medium effort. Treat it as a design to test on your own multi-file tasks, not a result to assume.