---
title: 'In the News: September 26, 2026 (Extra 3)'
description: 'A Google-led study finds automated agent-harness evolution overfits its benchmark, and tests a regularized version that still generalizes.'
canonical_url: 'https://darkfactory.dev/news/2026-09-26-extra-3'
markdown_url: 'https://darkfactory.dev/news/2026-09-26-extra-3.md'
collection: news
date_published: '2026-09-26T23:20:00-04:00'
date_modified: '2026-09-26T23:20:00-04:00'
---

# In the News: September 26, 2026 (Extra 3)


This edition covers one paper: a Google-led study showing that letting an LLM
automatically evolve an AI agent's own harness against a fixed set of
benchmark tasks tends to produce a harness that scores better on those tasks
and worse everywhere else, along with a regularized method that avoids that
trade.

## 1. A fix for harness self-improvement loops that quietly overfit their own benchmark

**[RRSI: Regularized Recursive Self-Improvement of Agent Harnesses](https://arxiv.org/abs/2609.24972)** · Peng Xia and 13 coauthors, Google Cloud AI Research, with Stanford University, Washington University in St. Louis, and UNC-Chapel Hill · arXiv, submitted September 21, revised September 23, 2026

Automating harness evolution, using an LLM to propose and keep edits to an agent's prompts, tools, and control flow based on scores from a fixed set of tasks, is a fast-growing shortcut for the manual work of harness engineering. The paper's own head-to-head comparison against four recent automated-evolution methods shows the risk of doing that without constraints: all four improve on the benchmark they were evolved against, but two of the four end up scoring worse than an untouched harness once moved to benchmarks they never saw. RRSI's fix caps how many edits a single round can bundle, screens out edits that encode benchmark-specific details, blocks any gain that falls inside measurement noise, and requires added inference cost to be paid for by a measured score improvement. Tested across eight benchmarks in coding, legal-document work, and engineering design, RRSI posts the smallest gain on the training benchmark of any of the five evolved harnesses compared head to head, but the only average gain on unseen benchmarks that clears the unevolved baseline by more than a single point, 43.6 against 39.7. AHE, the costliest rival, spends 58 percent more tokens per trial for 4.4 fewer points on that same held-out average. "No held-out split regresses anywhere," the authors write, "which is the failure a memorizing harness produces." The method also transfers to a smaller model that never took part in the search, raising its Terminal-Bench 2.1 accuracy from 11.2 to 14.6. Code, evaluation scripts, and the prompts given to the proposer and its critic are published on GitHub under an Apache 2.0 license.

**Why it matters:** A rising score from an automated harness-tuning loop is not proof the harness got better: two of this paper's own baselines improved on their training benchmark while getting worse on everything else. The checklist here, a shrinking edit budget, a noise floor before accepting a gain, and a cost check tied to measured improvement, gives anyone running that kind of loop a concrete way to catch the same failure before it ships.
