← In the News

An argument that benchmark pressure is why agents stopped asking questions

Why does Opus 5 feel worse to work with? · Mun Logadan · personal blog, August 14, 2026

Machine-readable Download Markdown

A short post, roughly 450 words, holding that Opus 5 is the more capable model on benchmarks and the worse one to work with, because Opus 4.7, Opus 4.8 and Fable "stop and ask questions if my intent was unclear" and "don't reinterpret or update my plans without asking." The proposed cause sits under a heading the author himself titled "Baseless speculation," and the label is his: "Selecting for models that do well on benchmarks (and indeed training for them or on RLVR tasks in general) inherently selects for models that make bold, usually-correct assumptions in the face of ambiguity. It penalizes models with a tendency to stop and ask for clarification or direction." No measurements are offered, and none are claimed.

Why it matters: The claim is unmeasured, so treat it as an argument rather than a finding to cite. It locates the cause of a harness problem in the evaluation arena, outside the model and harness. An agent that never stops to ask is exactly what lights-out operation requires and exactly what the author says he does not want. In his words, "Real life just isn't a benchmark."