I'm (mostly) picking models on speed now, not intelligence · Martin Alderson, cofounder, catchmetrics.io · martinalderson.com, 2 August 2026
A practitioner switches to picking models on speed, then names the ceiling on what speed buys
Alderson writes that for the first time he is choosing daily-driver models on tokens per second rather than raw capability, on the view that models around the Opus 4.6 level are good enough for most of his work. He proposes roughly 100 tok/s as the perceptual threshold, the output-side analogue of the 100ms interaction rule, and reports the spread for GLM 5.2 on OpenRouter as under 30 tok/s at the slow end to 129 tok/s at the fast end, with 109 tok/s from DeepInfra. He puts GLM 5.2 pricing at $0.42 and $1.32 per million tokens, which he calls 5% of the price of Opus. Those are figures he read off a marketplace, not measurements he ran.
He then limits his own argument. In an agent turn, model inference is only part of the wall clock, and tool calls and human oversight are the rest. "The 5x speedup on the model only buys you a 2x speedup on the turn, because the other 25 seconds didn't move." He labels this "rough numbers, but the shape holds," and it is introspective rather than instrumented. He expects HBM4 memory in Nvidia's Vera Rubin and AMD's MI400 parts to roughly double output token rates from bandwidth alone.
Why it matters: The Amdahl argument cuts against the headline. If your harness is bottlenecked on tool calls and review, a faster model returns less than its benchmark speedup, so the thing to instrument is your own turn breakdown before you re-pick models on tok/s. Alderson has not published that breakdown for his own setup, which is the measurement his argument most needs.