TDD inside the agent loop doesn't pay for itself, a small controlled study finds
Böckeler had Sonnet 4.6 build the same small, medium and large greenfield tasks with and without TDD instructions, twice each across five batches, then had Opus 4.8 blind-judge the resulting code and tests without knowing which workflow…