Aligned to whom? · Ryan Lopopolo, harness engineer, formerly OpenAI, now at Google Cloud · hyperbo.la, September 12, 2026
There is no such thing as an unhackable grader, a veteran harness engineer argues
Lopopolo starts with a practical limit: experts can catch an agent's mistakes in their own fields. Outside that expertise, in his examples double-entry accounting, finance, law, and operations, users rely entirely on the model's priors without knowing whether those priors are sound. He is blunt about his own field: "I am an expert software engineer and I am not happy (and never have been) with the default behaviors of the model when producing software."
His larger claim concerns grading itself. Lopopolo argues that every rubric, eval, and human rater that shaped a model's behavior brought its own blind spots, which the model's priors then inherited. "There is no such thing as an unhackable grader," he writes, because people disagree about what counts as a permissible shortcut. He connects this to a problem this feed has tracked before: models trained without a memory of past mistakes have no equivalent of "a fear of future regret." That helps explain why long-term coherence in agent-produced systems remains, in his words, "a very unsolved problem."
Why it matters: Before handing an agent unsupervised work outside your own expertise, ask who wrote the rubric it is optimizing against and whether that person's blind spots match yours.