A training method for self-evolving coding-agent harnesses that separates model-specific quirks from real harness bugs before rewriting anything, reporting a 1.84x training speedup and an 18.56 percent accuracy gain over prior methods on airline and retail agent benchmarks. This corpus already tracks several papers in this specific sub-field. The numbers refine an existing approach.
In the News: September 13, 2026
Ryan Lopopolo argues that no grader is unhackable, exposing the risk of trusting agents in domains where users cannot judge the work.
There is no such thing as an unhackable grader, a veteran harness engineer argues
Lopopolo starts with a practical limit: experts can catch an agent's mistakes in their own fields. Outside that expertise, in his examples double-entry accounting, finance, law, and operations, users rely entirely on the model's priors…
Read story →A researcher's case: your reproducibility habits are already agent context
Barba, a longtime reproducible-research advocate writing on agentic coding for the first time, maps five familiar software artifacts to the context a coding agent needs: a test suite, a clean commit history, a well-organized repository, a…
Read story →Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents
Garry Tan on X
Tan's quote-tweet was reported as saying that founders outside hardware are converging on building a domain-specific harness instead of remaining a system of record. The post had reached roughly 685,000 views, 3,400 likes, 269 reposts and 137 replies at 19 hours after posting. By the next morning, a separate trending topic tied to the same line had drawn more than 2,300 posts. The post itself and the underlying Demo Day observation have not been independently verified.