In the News: September 17, 2026
A UC Berkeley benchmark finds coding-agent harness choice barely changes success rates but can double the bill, weakening the case for a model default harness.
Midday edition
Machine-readable
Download Markdown
Story
Harness choice moves cost more than success rates
The researchers tested seven models across three harnesses, Claude Code, Codex CLI, and the minimal open-source harness Pi, on SWE-bench Lite and Terminal-Bench 2.0: 21 pairs, 30 sampled tasks repeated three times each. Claude Fable 5…
Read story →