You can date a period of this industry by the compound nouns it attaches to "token." Token oblivious. Token maxing. Token leaderboards. Token anxiety. Token minimizing. Each phrase is a fossil of a belief about what a token is worth, and if you walk the phrases in order, you get the whole history of how we have tried and mostly failed to reason about the cost of machine intelligence.
I want to walk them in order, because I think the story ends somewhere specific, and it is not where the current discourse is standing.
The unit everyone bills in and nobody thinks in
A token is a chunk of text, usually smaller than a word, that a model reads and writes. That is the mechanical definition, and it is the least interesting thing about it. The better framing comes from Nate B Jones, in a paid Substack post: "The unit of work is now the token. A token is a unit of purchased intelligence, fundamentally different from an instruction."
Two mechanical facts matter for everything that follows. First, there are three meters running, not one: input tokens, output tokens, and reasoning tokens, priced very differently. Output typically runs several times the input price, and reasoning is billed at output rates while remaining invisible in the response, which means a short answer can carry a long bill. Second, conversations compound. Because history is resent each turn, BCG notes that "the cumulative number of tokens billed across a session rises roughly with the square of the session's length. A session that feels twice as long can actually cost four times as much."
Era one: token oblivious
For the first couple of years, almost nobody thought about any of this, because flat subscriptions hid the meter. You paid your twenty dollars and hit a ceiling occasionally. The model companies subsidized usage to buy adoption, so the price of a token, from the seat you were sitting in, was zero. Nufar Gaspar, whose AI Daily Brief episode with Nathaniel Whittemore traces this same arc, calls it the all-inclusive era.
Whatever else you say about this period, it taught a generation of users that intelligence was free. Every era since has been a correction of that lesson.
Era two: token maxing
Then agents arrived, spend became real, and the correction overshot in the enthusiastic direction: spend became the metric.
The purest statement of the position is StrongDM's software factory principles: "If you haven't spent at least $1,000 on tokens today per human engineer, your software factory has room for improvement." Latent Space's write-up of Ryan Lopopolo, who runs a billion tokens a day, describes him as calling it borderline "negligent" to be running less. Meta tracked employee usage on an internal leaderboard, with monthly consumption in the tens of trillions of tokens. Engineers posted their bills publicly. As David Tepper put it to McKinsey, "You can see it in engineers posting $100,000 token bills on LinkedIn almost as a badge of honor. We've token-maxed like crazy, but there is no line that says what the business got for it."
It is easy to mock this era now, and I want to resist that, because the underlying logic was not stupid. Tokens were treated as fuel because for a certain kind of work they are. The constraint had visibly moved from compute to human attention, and a company that overspent while learning fast was plausibly better off than one that underspent while demanding ROI proof for every experiment. The maxers were wrong about the metric. They were not obviously wrong about the posture.
Era three: the bill arrives
The problem with usage as a badge is that usage compounds and badges do not.
Uber's 2026 AI coding budget was exhausted in four months. Per Mozilla's State of Open Source AI report, engineers were billing $500 to $2,000 a month before spend was capped at $1,500 per tool per employee. Meta went from leaderboard to constraint memo, and the press, with the industry's usual gift for symmetry, called it token minimizing. TechCrunch reported an unnamed company running up a nine-figure cloud bill with no usage limits in place.
So the pendulum swung, and it swung past prudence into something worse. The mood in most rooms now is what Gaspar calls token anxiety: practitioners feeling watched when they pick an expensive model, employees self-censoring to avoid costs, every prompt an implicit ROI conversation. Her line for it is the right one: the most expensive token is the one your best person is afraid to spend.
BCG's Return on AI work names both failure poles in one breath. Token maxing brings runaway bills, while minimizing or capping consumption starves high-return work and biases the organization toward narrow labor-substitution cases that are easy to measure but small in payoff.
That last clause is the important one. A cap does not make an organization efficient. It makes the organization do smaller things.
Both eras managed the same wrong number
Token maxing and token minimizing look like opposites, and they are the same mistake with the sign flipped. Both take total token spend, a numerator with no denominator, and manage it as if it were the thing itself. One maximizes the meter, one minimizes the meter, and neither can tell you whether any given token bought anything.
Tepper again, in the cleanest sentence written on this: "Tokens are not value. Tokens are the bill. The bill tells you what you spent. It does not tell you whether you should have spent it."
The correction is not a smarter attitude toward the meter. It is a denominator. Tepper's is cost per completed task. Sandeco Macedo's coding-agent survey sharpens it further with a discipline he notes practitioners almost never measure: "cost per accepted change, the tokens or money spent divided by the number of changes that survived verification." And for factory work I hold the long version: the total model, infrastructure, validation, retry, review, incident, and human-attention cost, divided by outcomes that are accepted and remain useful over time. The denominator charges you for the retries and the rework. Raw spend never does.
Put a denominator under the meter and the eras dissolve into ordinary engineering questions. High spend on work that survives verification is cheap. The Anthropic Economic Index found that compute tends to scale with the value of the work, which means an across-the-board minimization policy is, in expectation, a policy against your most valuable work. And unbounded spend on work that does not survive is not ambition, it is what Macedo calls a broken loop, one that "burns budget without producing approved changes" while looking busy.
The denominator also marks the real limit of maxing, which the leaderboards never found. Dex Horthy, on the review-agent arms race: "Of course, more review agents and more tokens do help -- they raise the floor, catching the dumb stuff. But they don't move the ceiling, because the ceiling is whatever we managed to teach the model in RL." Past the point where spend stops converting into accepted outcomes, more tokens are not more factory. They are exhaust.
Where this argument is weak
The era story is tidier than the reality. The eras overlap; there are teams token maxing today and teams that never left oblivion. The Meta and Uber numbers come from an internal memo reported by The Information and from press coverage, not audited disclosures, and I would not build on their precision. Tepper sells the measurement tooling his argument implies you need, which does not make him wrong but should be priced in. And the denominator I am advocating is much easier to define than to measure: "durable" only reveals itself over time, and attributing human attention to a task is somewhere between hard and self-deceiving. I argued previously that per-task cost attribution at the API boundary is the prerequisite, and I stand by that, but the numerator is still the easy half.
What this changes in my factory
I do not set a token budget. I set an acceptance bar and watch what it costs to clear it. When cost per accepted outcome rises, that is a signal about the loop, the harness, or the task shape, not a signal to use a cheaper model, although sometimes the answer is a cheaper model.
Spend that produces learning gets counted as production, deliberately. Gaspar calls these tokens that teach, and her point survives translation into my terms: an experiment that fails and tells you why is an accepted outcome with a long shelf life. It only looks like waste on a dashboard that has no denominator.
And I treat any spend I cannot attribute to an outcome as the highest-priority item on the bill, because it is the only category that is invisible to every other control. The $1,500 cap catches the engineer doing too much real work. It does not catch the compaction job running every thirty minutes against an empty session.
The pendulum has spent a year swinging between spend more and spend less because the meter was the only number in the room. The answer was never a better attitude toward the meter. Put a denominator in the room and the pendulum stops.