Sifry ran eight models through three ways of reviewing the same pull requests, a single engineered prompt and two open-source review harnesses (Compound Engineering and metareview), at three reasoning-effort levels. Every finding was…
In the News: September 21, 2026 (Extra 3)
A code-review harness beat a single prompt in a new public benchmark, and a purpose-built fact-checker outperformed general LLM reviewers at catching bad claims.
Extra edition
Machine-readable
Download Markdown
Story
A code-review harness beats a single prompt, and the cheaper model wins
Read story →
Story
A model built to check facts beat general reviewers at catching them
Tran tested five general-purpose models as reviewers of AI-written meeting summaries, hand-checking 35 claims against the source transcripts (23 supported, 12 not). Bigger reviewers caught more unsupported claims, up to 10 of 12, but the…