Benchmarking Fable 5, GPT-5.6 Sol, and Kimi K3 on SlopCodeBench · Dex Horthy · X, 4 August 2026
The first frontier-model results on SlopCodeBench: 33.3% strict pass, and the author will not run lights-off
Horthy published results for the new frontier on SlopCodeBench, the long-horizon coding benchmark from Gabe Orlanski's lab at UW Madison. Fable 5 and GPT-5.6 Sol tie at 33.3% strict pass, 10 of 30 checkpoints across 6 challenges, with Fable ahead on the isolated-pass tiebreaker at 16 to 14. Kimi K3 posts 26.7% on Modal and 23.3% on Baseten. In the prior run no model broke 25%. Slop-rule trip rates run 79% to 95% of all final code lines, with Horthy's own caveat that "some of the code quality measures are a bit over-aggressive."
Harnesses are deliberately mixed, Claude Code 2.1.219, Codex CLI 0.145.0 and OpenCode 1.18.0, so model and harness are not isolated. And on reliability: "I only did one run on each provider, so these results should not be read as statistically significant." His conclusion is the same one he held before the numbers moved: "The frontier is getting better, but I'm still not trusting them to run around lights off in my codebase." A follow-up experiment is announced, deterministic linters and alternating-model adversarial review inserted after each checkpoint, measured on strict pass rates.
The post could not be reopened independently. Every figure and quotation above comes from a contemporaneous end-to-end capture.
Why it matters: A 33.3% ceiling on a benchmark whose value is that it is unsaturated is a better number to plan against than a SWE-Bench percentage in the eighties. The announced follow-up will test whether harness intervention, rather than model choice, moves long-horizon completion.