← The Factory

Economics, capacity & factory FinOps

Optimize the whole queue for accepted, durable outcomes and scarce human attention, not tokens or lines of code.

medium confidence
Machine-readable Download Markdown

Confidence: medium. Evidence: strong gains in selected regimes; production self-report; weak standardized accounting. Last substantive change: 2026-08.

Model bills are only part of a factory's cost. Retries, verification, infrastructure, incidents, and scarce human attention all belong in the calculation.

The conclusion

Optimize the whole queue for accepted, durable outcomes and for scarce human attention. Parallelism is valuable only until validation, integration, infrastructure, or operational capacity becomes the bottleneck, at which point adding agents makes things worse. Measure model, harness, context overhead, retries, validation, and human attention together. The right denominator is cost per accepted, durable outcome, not tokens, lines, or raw task success.

How the thinking got here

Early framing counted tokens and lines of code. That gave way to per-task cost, then to cost per accepted change, then to attention economics, review backpressure, and downstream maintenance burden. The recurring surprise is that apparent technical deflation can coexist with rising verification and coordination cost, so a cheaper token does not always mean a cheaper outcome. A Databricks production account adds concrete cost infrastructure: a shared AI Gateway for model access, routing, budgets, client configuration, and session traces, with progressive spend gates and context-overhead reduction instead of relying only on hard per-user quotas.

Credible alternatives, and when each is right

Approach Right when
Frontier-model abundance quality dominates and budget is ample
Cheap-model cascades most work is easy, escalate the rest
Progressive spend gates and model downshifting preserving access while adding friction as spend rises
Hard quotas stopping runaway or unauthorized spend as a last resort
Value-based budgets tying spend to expected outcome value
Capacity-aware schedulers the bottleneck shifts across the queue

Where it fails and what we still don't know

Potential failures include Rémi Louf's Jevons-style hypothesis, where cheaper generation could increase total spend, and review queues that collapse under throughput. Databricks reports more than 30% lower average task cost from routing and almost 50% lower generated-token cost after harness and cache tuning, with no observed quality loss; these are useful first-party measurements, not independently replicated causal estimates. The corpus still does not show that falling token prices cause higher total spend, and standardized accounting across retries, review, maintenance, and incidents remains weak. Open questions include demand elasticity, full-cost accounting, marginal value curves, cost attribution, energy and carbon, and build-versus-buy.

What would change our mind

A standard full-cost accounting that includes validation, review, maintenance, and incidents would let factories compare themselves honestly rather than on token price.

Evidence and further reading