All articlesAI Trends

Model Routing Without Over-Engineering It

Not every task needs the strongest model. A simple routing rule beats a clever one you never maintain. How to split traffic by difficulty.

Sri Raman2 August 20267 min read
Model Routing Without Over-Engineering It

The strongest model is also the most expensive and slowest. Routing every request to it is the most common reason production LLM apps burn money for no quality gain. The fix is not a fancy router — it is a boring rule you will actually keep.

Classify difficulty, then route

Before calling the model, classify the request into cheap, medium, and hard with a fast, cheap classifier — often a small model or a few keywords. Route each tier to a model sized for it.

difficulty = classify(request)   // cheap call
model = { cheap: small, medium: mid, hard: strongest }[difficulty]
response = call(model, request)

Measure the mis-routes, not the average

Average cost goes down and average quality stays flat — which hides the cases where a hard query hit the small model and failed. Track the failure rate per tier and tighten the classifier until hard-tier-on-small-model is rare, not until the dashboard looks good.

Start with two tiers, not five

A two-tier rule (small for most, strong for flagged) is maintainable by a human. A five-tier learned router is a system you now operate. Begin with the rule you could write on a napkin and only add tiers when you have evidence the simpler one is leaving quality on the table.

The goal of routing is not to use the best model. It is to use the worst model that still passes the eval.
Share this article