Definition
Model routing selects which separately deployed model will handle a request or step. The router can use task characteristics, required capabilities, cost limits, and expected quality to make that selection. It can be a set of rules or a learned predictor.
For example, an application might send ordinary extraction requests to a smaller model and requests needing a larger context window to another. The router makes an allocation decision; the application remains responsible for checking the returned answer and enforcing the permitted actions.
Origin and attribution
Isaac Ong and coauthors introduced RouteLLM in June 2024. Their framework learns from preference data to choose between a stronger and weaker language model before generating a response. It is a documented implementation of model routing, with earlier selection methods acknowledged in the paper. The paper does not establish who first coined the broader term.
Model routing also includes explicit rules and catalogs with more than two models.
Scope and evidence limits
Routing among deployed models differs from selecting expert subnetworks inside a mixture-of-experts model. Both use selection, but the objects being selected and their operating costs differ.
RouteLLM's experiments found that routers trained only on Chatbot Arena preferences performed close to random on some out-of-distribution benchmarks. Adding relevant examples improved those results. A router's measured performance therefore needs the request distribution, candidate model versions, and evaluation procedure.
Operational significance
Include the router's own execution cost and latency in comparisons. Track which model was selected, why it was eligible, and what happened when it was unavailable or failed validation. Price changes, model updates, and changing task mixes can invalidate an earlier routing policy.
Test low-frequency tasks and costly errors alongside average quality. A rule that cheaply handles most requests can still select an unsuitable model for a consequential exception.
Distinguish it from nearby terms
- A model cascade inspects an earlier answer before deciding whether to escalate. Pre-generation routing selects without that answer.
- Mixture of experts routes inputs or tokens among subnetworks within a model. Model routing selects among deployed models.
- Calibration measures whether predicted probabilities match observed frequencies. It can support routing thresholds but does not define the routing policy.
- Orchestration coordinates a whole workflow. Model routing is one allocation decision within that workflow.
Check your understanding
A router reduces inference spend by choosing a small model for 90 percent of requests. Which measurements would reveal whether its overhead, rare mistakes, and changed task distribution erased the saving?