Model Routing: An Empirical Study
Empirical results on routing across foundation model families.
An empirical study of routing policies across open and proprietary models — accuracy, latency, and cost trade-offs on production workloads. Findings power default policies in AI Orchestrator.
Policies
Heuristic, learned, and hybrid routers.
Workloads
Chat, extraction, code, and vision.
Metrics
Quality-adjusted latency and cost per task.
Recommendations
Practical defaults for enterprise stacks.
Setup
We evaluate routers on anonymized enterprise traffic spanning chat assistance, document extraction, code generation, and vision Q&A. Models range from small open weights to frontier APIs.
Results
Hybrid routers — heuristics for obvious cases, learned classifiers for the ambiguous middle, cascades for high-stakes tasks — dominate pure strategies on quality-adjusted cost. We publish decision charts practitioners can adopt.
Caveats
Router quality depends on fresh labels. We discuss how often to retrain, how to detect when a new model invalidates old policies, and how to keep humans in the evaluation loop.
Discuss this work with our labs.
