← In the News

Confidence scores are becoming an excuse, not a guardrail

The Normalization of Inexplicable Failures · Patrick Xia, i hate the future · Sep 27, 2026

Machine-readable Download Markdown

Xia argues that teams shipping AI-driven features often treat a confidence score as a stand-in for real verification instead of a calibrated signal. Writing about a product he calls Jev, built on TypeSafe AI's classification and scoring primitives, he points to the vendor's own documentation, which sets example thresholds (0.5 to act, 0.9 for high-risk actions) while cautioning that "the correct threshold values depend on your domain," a caveat he says most builders skip past. "At best, people use confidence scores in a cargo cult manner. At worst, people use them as an excuse for why the API call failed," he writes. The piece had drawn 81 points and 14 comments on Hacker News by 1:15 p.m. ET Sunday, about 1.6 hours after it posted.

Why it matters: Xia's real complaint is about accountability, not confidence math. A team that skips building an eval and a ground-truth pipeline loses the ability to trace a failure back to a cause, even though the same AI-accelerated tooling that shipped the shaky feature could just as easily have produced the eval. "My fear is that 'sometimes it just sucks' is going to be more and more the accepted endpoint of investigations," he writes, and that is exactly the failure mode this lane keeps warning about: a guardrail that only looks like verification.