In the News: September 8, 2026
A 26-condition eval finds that instructing coding agents to use a specific test technique, including TDD and formal methods, mostly does not improve correctness.
Extra edition
Machine-readable
Download Markdown
Story
Telling coding agents to use TDD made results worse in a 26-condition eval
Luu re-ran his earlier Zstd-in-Rust agent eval with Codex (GPT-5.6 Sol). He tested 26 different instructions, ranging from "use test-driven development" to formal tools such as TLA+, Lean 4, Kani, Verus, and SMT solvers, along with four…
Read story →