Introduction
An LLM router decides which model handles each request, and in 2026 it is the single biggest lever on what an AI system costs to run. Most teams pick one model at prototype time and send everything to it: the trivial classification, the long reasoning step and the meta description all go to the same flagship at the same price. The bill grows with usage, and nobody can say which step is responsible for it.
This guide is for CTOs and engineering leads who already have an AI feature or agent in production and want to stop paying flagship prices for work a cheaper model does equally well. I will explain what LLM routers are, how a router differs from an AI gateway, the main routing strategies, and how we built and tested routing in our own production systems.
You will also find an honest overview of the top LLM routers and gateways on the market, and a section on when building a router is not worth it. Every claim about how we route comes from our own code, and where we have no measured number, I say so.
What Is an LLM Router?
An LLM router is a layer between your application and the models it calls. For each request, it chooses a model (and sometimes a provider) based on rules, a classifier or learned preferences, then forwards the request and returns the answer. In other words, model routing turns "which of the available LLMs should answer this?" into an explicit, testable decision. The application asks for "an answer"; the router decides who produces it.
What LLM Routers Do in Practice
LLM routers make three kinds of decisions:
- Model selection: which model should handle this request, given its difficulty, its type and the cost you are willing to pay.
- Provider selection: which inference provider should serve that model, given price, speed and availability.
- Failover: what to try next when the chosen model or provider fails.
Not every router does all three. Some only pick a model, some only balance traffic across providers, and many teams build the first themselves and rent the other two.
Router vs Load Balancing vs Fallback
These terms get mixed up, so it is worth separating them. Load balancing spreads identical requests across several endpoints that serve the same model, to protect throughput and availability. A fallback is a backup used only when the first choice fails. Routing is the decision about which model a request deserves in the first place.
A mature routing layer usually combines all three, but they solve different problems. If your bill is the problem, you need routing. If outages are the problem, you need failover. If rate limits are the problem, you need load balancing.
Why a System With One Model Fails
A system one model serves end to end is simple to build and easy to reason about, which is why most prototypes look like that. It breaks down at scale for two reasons. First, cost: most requests in a typical AI product are short and easy, and paying flagship prices for them is waste. Second, resilience: when that one model or provider is slow or down, the whole product is down with it.
Routing addresses both. It lets cheaper models handle easy queries, keeps strong models for hard work, and gives every request somewhere else to go when the first choice fails.
LLM Router vs AI Gateway: What Is the Difference?
This is one of the most common questions about LLM routers, and the honest answer is that the categories overlap. An AI gateway (also called an LLM gateway) is infrastructure that sits in front of model providers and handles the plumbing: one API for many providers, authentication, logging, rate limiting, caching, retries and cost analytics. A router is the decision logic that picks the model.
What an AI Gateway Does
According to Cloudflare's AI Gateway documentation, its gateway offers analytics on requests, tokens and cost, logging, response caching, rate limiting, and request retry with model fallback. That list is typical: a gateway is about control and visibility over LLM traffic, and the routing it offers is usually rule-based or failure-based rather than a decision about which model a request deserves.
Where the Two Overlap
The lines blur because most gateways now include some routing, and some routers are delivered as gateways. OpenRouter is a good example: it is a unified LLM API across hundreds of models, which makes it a gateway, and it also routes each request across the inference providers that serve a model, which makes it a provider router. What it does not decide by default is which model your request should use; that stays with you unless you opt into its automatic router.
How We Combine Them
In our own systems, OpenRouter is the gateway: one API key, one billing account and one request format for every model we use. On top of it, our code makes the model decisions, either through configuration by role or through a classifier, and sets provider preferences per request. On our AI fundraising platform, we also send document extraction and aggregation calls through Cloudflare AI Gateway, while the chat routers call OpenRouter directly.
The practical rule we follow: rent the gateway, own the routing policy. Gateways are commodity infrastructure, but the decision about which model a request needs depends on your tasks, your quality bar and your evaluation data, and that knowledge belongs in your code.
Routing Strategies: Five Ways to Pick a Model
LLM routing is not one technique. There are five common routing strategies, and each trades decision quality against cost and latency. They are not exclusive; most production systems use two or three of them at different layers.
Static Routing by Role
The simplest strategy assigns a model to each step of the system once, in configuration. The AI editor that runs our own content operations routes this way, with defaults defined in one model configuration file:
- The orchestrator that runs long tool-calling loops uses
z-ai/glm-5.2. - Section writing, chart generation and image prompts use
google/gemini-3.8-flash. - SEO verification and data analysis, which need strong reasoning, use
openai/gpt-6-sol. - Meta titles, keyword optimisation and vision tasks use
openai/gpt-6-luna, described in the config as "fast and cost-effective". - Web search goes to
perplexity/sonar.
Each brand can override any role from the admin panel without a deploy. The resolution order is per-call override, then the brand's saved setting, then the hardcoded default. That design matters more than the model names: you can move one role to a cheaper model, watch quality, and move it back in minutes.
LLM-Based Routing With a Classifier
What is LLM-based routing and how does it work? A small, fast model reads the incoming request and classifies it into one of a fixed set of categories; each category maps to a downstream handler and model. The classifier is cheap because it outputs a label rather than an answer, and the expensive model only runs when the category needs it.
This is the strategy we use when the right model depends on what the user has said, which static roles cannot know in advance. The next section walks through our production version in detail.
Semantic Routing With Embeddings
A semantic router skips the classifier model entirely. It embeds the incoming request and compares it with example utterances for each route, picking the closest match. Aurelio Labs' semantic-router library is the best-known open source implementation, and its pitch is speed: a vector comparison takes milliseconds, while an LLM classifier takes a model call.
We have not needed semantic routing in production. Our routing decisions depend on conversation state (which question we asked, what the user already answered), and an LLM classifier reads that state more reliably than a similarity score. For high-volume intent detection over short, independent messages, a semantic router is a reasonable choice.
Learned Routers Trained on Preference Data
A learned router predicts, for each query, whether a cheaper model will answer with the same response quality as a stronger one. RouteLLM is the reference framework here: the RouteLLM paper trains routers on human preference data with data augmentation and reports cost reductions of over two times in certain cases without compromising response quality, with routers that kept working when paired with different models than those used in training.
Learned routers are attractive when you have high traffic of varied queries and a clear pair of models, one strong and one cheap. They need training data and benchmarks that reflect your own traffic, which most teams do not have at the start, so we treat them as an optimisation for later, not a starting point.
Provider Routing: Same Model, Different Inference Providers
Open-weight models are often served by several inference providers at different prices and speeds, so routing also happens one level down. OpenRouter's provider routing documentation explains that by default it load-balances across providers, weighting each by the inverse square of its price, and skipping providers with recent outages. You can override that with an explicit provider order, restrict the provider list, or sort by price, throughput or latency.
We use these controls deliberately. On our fundraising platform, the intake router runs openai/gpt-oss-120b restricted to fast providers (Groq, Baseten, Novita and Cerebras) sorted by throughput, and a document extraction step uses the :nitro variant of the same model with an explicit provider order. In an internal project-analysis tool, each model is pinned to a provider and quantisation in configuration, so results stay comparable between runs.
How We Built an LLM Router in Production
Our most complete router runs inside an AI fundraising platform that guides founders through a long intake conversation and turns their answers into investor materials. Every user message has to be interpreted in context: is it an answer, a partial answer, a request for help, or an attempt to skip the question?
Stage One: Classify the Message
The intake router reads the current question and the user's message and returns a JSON verdict with one of six actions:
| Action | Meaning | What happens next |
|---|---|---|
| USERRESPONSEVALID | The user answered the question | The answer is parsed and stored |
| PARTIAL | The user answered part of it | The partial answer is parsed and stored |
| INVALID | The reply does not answer the question | The user is re-prompted with a hint |
| SKIP | The user clearly wants to skip | The field is marked as not provided |
| RESPOND | A short reply is enough | The router's own hint is sent, no further model call |
| EXPLAIN | The user needs an explanation | A second router picks the explanation model |
The RESPOND path is easy to miss and important for cost. When the router can answer with its own short hint, the request ends there, and no second model is called at all.
Stage Two: Pick the Downstream Model
When the first stage returns EXPLAIN, a second router decides how much model the explanation needs: REGULAR, DETAILED or WEBSEARCH. Its prompt tells it to prefer REGULAR unless a more detailed explanation is explicitly requested, and to choose WEBSEARCH only when the answer needs access to the web.
Each mode maps to a different model tier. REGULAR goes to google/gemini-2.5-flash, DETAILED goes to Claude Haiku 4.5 with Claude Sonnet 4.5 as backup, and WEB_SEARCH goes to Claude Sonnet 4.5 with search enabled. The expensive model with web access only runs when a classifier has decided the request needs it.
A Fast, Cheap Model for the Router Itself
A router adds a model call to every request, so the router has to be cheap and fast or it eats its own savings. Our intake router runs on an open-weight model, openai/gpt-oss-120b, on throughput-sorted providers, and the explanation and edit routers run on google/gemini-2.5-flash. The router's job is a short classification, so it does not need a frontier model.
Structured Output Keeps the Router Honest
A router that returns free text is a router you have to parse with regexes. Our intake router returns JSON enforced by a schema at the request level, validated with Zod on the way back, and repaired with a JSON-repair step if the model's output is slightly malformed. An invalid verdict never reaches the business logic; it either gets repaired or triggers the next fallback.
Testing the Router With Promptfoo
Routers fail quietly: a misclassified message still produces a reply, it is the wrong kind of reply. We test the routing prompt with promptfoo, generating test cases from combinations of real question types and user reply patterns, more than 1,200 cases across categories such as guidance requests, skips, requests for examples and "I don't know" answers. Each case checks the verdict twice, once with an llm-rubric judge and once with a deterministic check on the returned JSON.
The suite is organised by category, so a prompt change can be tested against the categories it affects without running every case, which keeps evaluation fast enough to use during development. Our guide to how we run LLM evaluation across models before shipping covers the tooling in depth.
Measure Cost Before You Route
A router without cost data is guesswork. You cannot know whether routing saved money, or which route is expensive, unless every LLM call is priced in dollars and rolled up to the task. This is the first thing we build into any agent, before any routing logic.
The Token Types That Make Up the Bill
A single LLM call is priced on up to six token types, and treating them as "input" and "output" is how cost dashboards end up wrong by a factor of two:
| Token type | How it is billed | Why it matters |
|---|---|---|
| Input (uncached) | Base input rate | The prompt, history and tool results sent fresh |
| Output | Output rate, usually several times input | The visible answer |
| Reasoning | Output rate unless the provider says otherwise | Thinking tokens you pay for but never show |
| Cache read | A fraction of the input rate | Prefix tokens served from cache |
| Cache write (5 minute) | A premium over the input rate | Creating a short-lived cache entry |
| Cache write (1 hour) | A higher premium | Creating a long-lived cache entry |
In the shared inference library we use across our AI workspace and a client document portal, one calculateCost() function prices all six. It subtracts reasoning tokens from output before pricing, warns when reasoning exceeds output, and reports how much the cache saved compared with the base rate.
Where the Prices Come From
Our pricing module pulls live per-model prices from OpenRouter's model list and refreshes them every ten minutes, with preset prices as a fallback. When the provider reports a cost on the response itself, we trust that figure verbatim. Missing rates are handled conservatively: an unknown cache-read rate is charged at zero with a warning, an unknown cache-write rate at the full input rate with a warning.
Roll It Up to the Task
Our agent runtime merges every subagent's tokens and cost into the parent message and keeps a running total per conversation, stored on each chat and shown in a developer view with a per-message breakdown. With that in place, you can compare routes on real data: cost per task by route, quality per route from your evaluation suite, and latency per route from your traces. Without those three numbers, LLM routing decisions are opinions.
Failover: What Happens When a Model Fails
Routing to cheaper models and smaller providers raises a fair question: what happens when they fail? Failover is the part of a routing layer that protects availability, and getting it wrong doubles your cost.
Classify the Error Before Retrying
Our agent runtime classifies every error before deciding what to do. Rate limits, overloads, server errors, network failures and stream stalls count as transient, and authentication, quota and model-not-found errors count as configuration problems; both move the request to the next model in an ordered fallback list.
Request errors never fall back. A context-length overflow, a content-policy refusal or a malformed request will fail the same way on the next model, so retrying them only pays twice for the same failure. Truncated output triggers a continuation call instead of a full retry, which keeps the tokens already generated.
Fallback Chains in Practice
On the fundraising platform, every routing stage has its own chain. The explanation router falls back from google/gemini-2.5-flash to anthropic/claude-3-haiku and then to openai/gpt-4o-mini, and document edits fall back from Claude Sonnet to Gemini Pro. In each chain, the backups come from different vendors than the primary, so one vendor's outage does not take out the whole chain.
Keep Data Rules on Every Attempt
A fallback is a new request, and it must carry the same data rules as the first attempt. In our shared library, zero data retention and a "deny" data-collection setting are merged into the OpenRouter settings of every attempt by default, rather than written into each attempt's definition. That way a fallback added in a hurry cannot silently send data to a provider that stores it.
Choosing the Right Model on Accuracy per Dollar
Model routing only saves money if the cheaper route is good enough, so quality has to be measured per route, not assumed. Leaderboards rank models on general benchmarks; your task has its own fields and failure modes, and the only model performance ranking that matters is the one you measure on it.
A Readout From a Document Extraction Project
On an invoice and contract processing platform we built for a fintech client, we compared four models on hand-labelled documents (10 invoices and 9 contracts) and checked reliability across 85 documents:
| Model | Invoice F1 | Contract F1 | Cost per document |
|---|---|---|---|
| Claude Sonnet 4.6 | 98.4% | 91.0% | $0.0322 |
| Claude Haiku 4.5 | 96.4% | 84.8% | $0.0187 |
| Amazon Nova 2 Lite | 97.5% | 83.3% | $0.0076 |
| gpt-oss-20b | 95.9% | 79.6% | $0.0083 |
Financial fields scored 100% on the leading models; the most expensive model bought better contract metadata, such as titles and start dates, not better financial accuracy. So the production defaults route by document type: invoices to Nova 2 Lite, contracts to Sonnet. That is routing at its simplest, a rule based on input type, and it kept invoices from costing about four times more for no measurable gain.
Routing Inside a Client Backend
On a residential property platform we build the backend for, the same idea lives in plain YAML. A deliberately cheaper "lite" model handles short translations, with cost reduction listed as a client requirement, while a stronger "pro" model is reserved for extracting the scope of a contract, which needs long context and runs once per contract. Expensive models are fine on rare, high-stakes steps; the steps that run on every request are where the cheap model has to earn its place.
Send Fewer Tokens Before You Route
The largest saving on the fintech project came before any routing. An early version of the extraction flow sent around 95,000 tokens per document through our own configuration mistakes; the corrected flow sends roughly 3,000 in and 300 out, which put the realistic cost at $0.006 to $0.015 per document. Check what each call carries before you tune which model receives it; our article on how we decide what an AI agent sees covers that side.
Caching is the other half of the same story: it discounts the part of each prompt that repeats, and it works best when the routing layer keeps requests for one conversation on the same model and provider. We cover provider-specific rules and our own measurements in our practical guide to prompt caching.
Top LLM Routers and AI Gateways in 2026
Which LLM router is the best? There is no single answer, because the tools solve different layers of the problem. Here is an overview of the platforms we see most often, what each is for, and where it fits. We use OpenRouter and Cloudflare AI Gateway in production; the others are described from their own documentation.
OpenRouter
OpenRouter is a unified LLM API and gateway across hundreds of models, with provider routing (price-weighted load balancing by default, or explicit order, price, throughput or latency sorting) and per-request data controls such as zero data retention. It also offers an Auto Router that classifies each prompt into a task type and picks a model based on what the wider market spends on for that task, within a cost tier you set. It is our default gateway because it keeps one API, one bill and one set of data rules across every model we use.
LiteLLM
LiteLLM is an MIT-licensed open source library and proxy that gives many providers an OpenAI-compatible interface. LiteLLM's routing documentation describes load balancing strategies including simple shuffle by rate limits, least busy, lowest latency, usage-based and cost-based routing, plus fallbacks, retries and cooldowns for failing deployments. It suits teams that want to self-host the gateway layer.
Portkey
Portkey's AI Gateway is an MIT-licensed open source gateway with fallbacks, load balancing, retries with exponential backoff and conditional, rules-based routing, according to the Portkey gateway repository. It is aimed at teams that want gateway features with an observability platform around them.
Cloudflare AI Gateway
Cloudflare AI Gateway sits in front of providers and adds analytics, logging, caching, rate limiting, and retries with model fallback. It fits teams already on Cloudflare Workers, which is where we use it: document extraction workers on our fundraising platform reach models through it.
RouteLLM
RouteLLM is an open source framework for learned routing between a strong and a weak model, trained on preference data. It is the closest thing to "intelligent model routing" as a reusable component, and the right place to start if you want a router that decides difficulty per query rather than by rules.
Semantic Router
Aurelio Labs' semantic-router is an MIT-licensed library that routes by embedding similarity to example utterances, without calling an LLM for the decision. It is the fastest option for intent detection when routes are stable and messages are short.
So Which One Should You Start With?
Start with a gateway that gives you one API, one bill and transparent pricing across providers, route by role in your own configuration, and add a classifier only where the right model depends on the message. Learned routers and semantic routing are optimisations for high, varied traffic. The best LLM router is the simplest one your evaluation data says is good enough.
When Not to Build a Router
Not every system needs routing logic. If your product handles a few hundred queries a month, or every request genuinely needs the same strong model, a router adds a model call, latency and a new failure mode for savings you will not notice.
Three cases where we would not build one yet:
- Low volume: price every call from day one, because that is cheap, and add routing when the numbers justify it.
- No evaluation set: without tests on your own cases, you cannot tell whether the cheaper route is good enough, and routing becomes a quality gamble.
- Uniformly hard tasks: if every request needs long reasoning, there is no easy traffic to route away from the expensive model.
The order that works for us is measure, trim the context, route by role, add a classifier where the message decides the model, and only then consider learned routers or owning a model. When one narrow task carries enough volume, owning a fine-tuned small model can beat any router, which we cover in our guide to running a self-hosted LLM for business and our comparison of small language models vs large ones.
Matt Sadowski
CEO of Mobile Reality
Cut What Your AI Agent Costs per Task
We instrument AI agents to show what every task costs, then lower the bill without lowering quality.
Per-call USD pricing rolled up per task and per conversation.
Context trimming, so each call carries only what the step needs.
Role-based model routing and prompt caching designed for your providers.
Fallbacks that never retry errors that will fail again.
AI automation measured before and after every change.
Conclusion
An LLM router is the decision layer that sends each request to the model it needs, and it pays off when it is built on measured cost and measured quality rather than on leaderboard ranks. The pattern we use in production is simple to describe and takes discipline to keep.
- Separate the gateway from the routing policy: rent one API for every provider, and keep the decision about which model a request needs in your own code.
- Start with static routing by role in configuration, and add an LLM classifier where the right model depends on the message.
- Make the router itself cheap and structured: a fast model, a strict JSON schema and its own fallback chain.
- Price every LLM call by token type and roll it up to the task, so you can compare routes on real cost.
- Classify errors before failing over, and carry the same data rules into every fallback attempt.
If your AI bill grows faster than your usage, start by measuring what each task costs. Our team can instrument an existing system, show you where the money goes and design the routing around it, through our AI automation agency services.
Frequently Asked Questions
What are LLM routers?
LLM routers are a layer between your application and the models it calls. For each request, the router chooses which model, and sometimes which inference provider, should handle it, based on rules, a classifier or learned preferences. The goal is to send easy requests to cheaper models and keep strong models for the work that needs them, with a fallback when a model fails.
What is the difference between an LLM router and an AI gateway?
An AI gateway is infrastructure in front of model providers: one API, authentication, logging, rate limits, caching, retries and cost analytics. An LLM router is the decision logic that picks which model a request deserves. Many gateways now include some routing, so the two overlap; we rent the gateway and keep the routing policy in our own code.
What is LLM-based routing and how does it work?
A small, fast model reads each incoming request and classifies it into a fixed set of categories, usually returning JSON. Each category maps to a downstream handler and model, so the expensive model only runs when the category needs it. In our production router, one category even lets the router answer directly, with no second model call.
Which LLM router is the best?
There is no single best router, because tools solve different layers. OpenRouter, LiteLLM, Portkey and Cloudflare AI Gateway handle the gateway layer, while RouteLLM and semantic-router focus on the routing decision. Start with a gateway plus routing by role in your own configuration, and add a classifier or learned router only when your evaluation data shows it pays off.
More on AI Cost, Small Models and Evaluation
What an AI system costs, which model it runs on, and how you prove it still works are one decision, not three. These articles cover the build and run cost, model choice and the evaluation that makes switching safe:
- Prompt Caching: A Practical Guide for Claude, OpenAI, Gemini
- Self-Hosted LLM Guide: Hardware, Tools and When It Pays Off
- SLM vs LLM: Key Differences and When Small Models Win
- Fine-Tuning LLMs: A Guide From Our Gemma 4 Case Study
- AI Development Cost in 2026: What Drives It and Real Ranges
- LLM Evaluation in Practice: How We Test Across 32 Models
- Context Is King: How to Serve It With Context Engineering
- Structured LLM Output Without JSON Schemas | MDMA
- Build AI Agents with 75+ Deployments Cutting Costs 60% in 2026
Want to know whether a smaller or self-hosted model could handle part of your workload? Our custom AI model development team starts with an evaluation on your own cases.
