The situation
A product team runs a customer-facing assistant on Amazon Bedrock. Every request goes to one large capable model. The bill has grown faster than usage: the model is priced for its hardest job and asked to do its easiest one thousands of times a day. The traffic is a mix. A large share is simple: classify an incoming message into one of a handful of intents, pull an order number out of a sentence, answer a factual question the knowledge base already contains. A smaller share is genuinely hard: reconcile a multi-part complaint, work through a returns policy with three conditions, plan a sequence of steps and explain the reasoning.
When the team samples the logs, the split is roughly eighty-twenty. Four in five requests are the kind a small model answers correctly and in a fraction of the time, at a fraction of the per-token price. One in five actually exercises the large model’s reasoning. Running everything through the flagship means the eighty per cent subsidises the twenty, in cost and in the extra latency the big model adds to requests that never needed it.
The obvious move is to send easy requests to a cheap small model and hard ones to a large capable model. The obvious risk is getting the split wrong. Route a hard request to the weak model and you get an answer that is fast, cheap, and wrong, phrased exactly like a right one. The decision is how to classify each request, how much a misroute costs, and whether to build the router or let Bedrock run it.
What actually matters
Model size is a cost lever, not a quality dial you turn up for safety. A bigger model is not uniformly better at everything; it is better at the tasks that need what it has, which is depth of reasoning and breadth of knowledge. On a single-step classification, a well-chosen small model matches it and answers quicker. So the question is never “is the big model better” in the abstract, it is “does this particular request need what the big model has”. Where the answer is no, you are paying for capability the request will not use.
The shape of the workload decides how much routing can save. If almost every request is hard, there is little to route away and the router saves nothing. If the traffic is genuinely mixed, with a large easy majority and a smaller hard tail, routing moves most of the flagship’s volume onto the cheap model. Quality on the hard tail stays where it was. The eighty-twenty split is the case routing is built for. A workload that is uniformly hard, or uniformly trivial, needs the one model that fits and no router.
The cost of a misroute is asymmetric and it runs one way harder than the other. Send an easy request to the big model and you overpay by a few tokens and a little latency; the answer is still correct. Send a hard request to the small model and you can get a wrong answer that reads as authoritative. It reaches the customer, and costs far more than the tokens you saved. Set the escalation threshold with that asymmetry in mind. Sending a borderline-easy request to the big model costs less than sending a hard one to the small model. When in doubt, route up.
That means the thing you actually have to measure is quality per route, not quality in aggregate. An overall accuracy number hides the failure that matters, because the misroutes are a minority inside a mostly-correct stream. You want to know how often the small model was handed something it got wrong. That means sampling the requests that took the cheap path and checking them against what the capable model produces. Without per-route measurement you cannot tell a healthy router from one that is degrading answers to save money.
And routing is not the only cost lever, so it should not be the only one you pull. A cache in front of the whole thing removes the repeated identical and near-identical requests before either model runs; trimming a bloated prompt cuts the per-call cost on both paths. Routing decides which model a request reaches; caching and prompt trimming decide whether it needs a model at all and how much it costs when it does. They compound, and the cheapest request is the one the cache serves.
What we’ll filter on
- Workload variety, is the traffic a genuine mix of easy and hard, or mostly one kind?
- Misroute cost, how bad is a wrong answer from the weak model, and how asymmetric is it against overpaying on the strong one?
- Classification signal, can the request’s difficulty be told from cheap signals (task type, length, a small classifier) or does it need the model to look?
- Measurability, can quality be measured per route, so a degrading cheap path is visible?
- Build versus managed, does a managed router within a model family fit, or does the split need custom logic across families?
- Stacking, does routing sit alongside caching and prompt trimming rather than replacing them?
The landscape
A rules or heuristic router classifies each request from cheap signals before any model runs. Route by task type when the caller already knows what it is asking for: a classification endpoint to the small model, a reasoning endpoint to the large one. Route by input length as a rough proxy. Short inputs are more often simple and very long ones more often need the bigger context, though length is a weak signal on its own. Route on the output of a tiny classifier, a small cheap model or a lightweight text classifier whose only job is to label the request easy or hard. Heuristic routing is transparent, easy to reason about, and free of an extra model call when it keys on task type or length. Its weakness is that a hand-written rule cannot detect difficulty the surface of the request does not show.
API-based model cascading runs the cheap model first and promotes to the capable one when the cheap answer looks weak. The small model attempts every request; if it signals low confidence, or a validator finds the answer malformed or failing a check, the request is retried on the large model. This adapts to difficulty the request’s surface does not reveal, because the cheap model has actually attempted the task. The cost is that hard requests pay twice, once for the failed cheap attempt and again for the capable retry. Cascading works when the easy majority is large enough that the doubled tail stays cheap overall. It also depends on a reliable signal that the cheap answer was inadequate.
Amazon Bedrock Intelligent Prompt Routing is the managed option, generally available since April 2025. For each prompt it predicts the response quality of the candidate models and forwards the request to the one that gives the best quality for the cost. A router holds exactly two models from the same family, and one of the two is nominated as the fallback. You either take a Bedrock-provided default router or configure your own pair from the Anthropic, Meta or Amazon Nova families. The dial you set is the response quality difference, measured against the fallback. At ten per cent, the request stays with the fallback unless the other model’s predicted response is ten per cent better. Requests reach it through the Converse and InvokeModel APIs, addressed to the router ARN rather than a model id, and the Converse response names the model that served it in trace.promptRouter.invokedModelId. AWS states that routing can cut costs by up to thirty per cent without compromising accuracy. Routing itself is charged at USD$1 per 1,000 requests, on top of the tokens the serving model bills for. Two constraints bound it: the pair sits inside one family, and the routing is tuned for English prompts.
Fanning out and aggregating inverts the trade. Where the routers above pick one model per request, an ensemble uses several in concert. Send the same request to two or more of them in parallel, and combine what comes back with your own aggregation logic. A majority vote settles a classification, and best-of-N returns whichever candidate a judge model scores highest. A merge takes the field each model is strongest on: one for the sentiment, another for the extracted entities, a third for the policy citation. The cost shape is plain. N models means N times the token spend and the latency of the slowest of them, so fanning out spends more rather than less. The extra spend is worth it where being wrong is expensive: a compliance classification that triggers a filing, a triage flag that sets who is seen first. It is wasted on the intent labels that make up most of this traffic. Ensembles and routers are both ways of selecting a model. One spends extra to raise confidence on a narrow slice of traffic; the other spends less to hold quality steady across a wide one.
Intelligent model routing systems have to be implemented somewhere, and where the decision runs is separate from what the decision says. Static routing configurations in application code are the plainest, a map from task type to model id that ships with the release and changes when you deploy. Step Functions for dynamic content-based routing puts the choice in a Choice state. The state machine reads a field off the request and branches to a different model on each arm, which fits when the surrounding workflow already runs there. API Gateway moves the logic to the edge, rewriting the target model from a header or a path segment before the request reaches your code. That suits a platform whose callers pick their own tier. A metric-driven router reads recent latency, error rates or throttling counts before choosing, so a model that has started timing out sheds traffic without anyone deploying anything. Each home changes who can alter a route and how fast: a config in code needs a release, a Choice state needs a state-machine update, an edge transformation needs neither.
Underneath all of them is the same measurement loop. Whichever router you run, sample the cheap path, check its answers against the capable model, and tune the threshold from what you find rather than from a guess.
Evaluation
Side by side
| Approach | Sees hidden difficulty | Extra call on easy path | Hard requests pay twice | Managed | Tuning surface |
|---|---|---|---|---|---|
| Rules by task type | ✗ | ✗ | ✗ | ✗ | Route table |
| Rules by input length | ✗ | ✗ | ✗ | ✗ | Length cut-off |
| Small classifier router | Partly | ✓ (tiny) | ✗ | ✗ | Classifier and threshold |
| API-based model cascading | ✓ | ✗ | ✓ | ✗ | Confidence and validator |
| Fan-out ensemble | ✓ | ✓ (N calls) | ✓ (all of them) | ✗ | Model set, aggregation rule |
| Bedrock Intelligent Prompt Routing | Predicted | ✗ | ✗ | ✓ | Quality difference, fallback |
The routed path
The solution
Start with the workload before touching a router, because the whole case rests on the split. Sample real traffic, sort it into easy and hard by hand, and get the ratio. A genuine eighty-twenty is the sweet spot: a large easy majority to move off the flagship and a real hard tail to protect. If the sample comes back mostly hard, routing saves little and adds a moving part for nothing, and the honest answer is to keep the one model that fits. If it comes back almost all trivial, drop to the small model outright and skip the router. Only a mixed stream justifies a router at all.
For a mixed stream where the difficulty is legible from the request itself, a rules or classifier router is the least machinery. When each endpoint already knows its task, a route table by task type is transparent and adds no extra model call. The intent classifier and the extraction job point at the small model, the reasoning endpoint at the large one. Where task type is not enough, a tiny classifier that labels the request easy or hard gives a cheap signal without running the expensive model first. Length can supplement this as a coarse tiebreak, but do not lean on it alone. A short input can still be a hard reasoning problem, and a long one can be a simple extraction from a wall of text.
When difficulty is not visible on the surface, API-based model cascading is worth it, with the escalation threshold set against the asymmetry of a misroute. The cheap model attempts everything; a low-confidence signal or a failed validation promotes the request to the capable model. A wrong cheap answer costs more than an unnecessary escalation, so tune the threshold to escalate readily. Sending some easy requests up beats letting hard ones through on the cheap path. The trade is that promoted requests pay for both models. The pattern depends on an easy majority large enough to keep the doubled tail cheap, and on a promotion signal you trust.
Amazon Bedrock Intelligent Prompt Routing is the pick when a single model family spans the capability range and you would rather not build and maintain the router. Point the request at a router instead of a named model, and Bedrock predicts which of the two members will answer better. The request stays with the fallback unless the other clears the response quality difference you set. Routing is billed at USD$1 per 1,000 requests, on top of the tokens the serving model bills for. A tenth of a cent is immaterial against a flagship call, and worth checking against the per-call cost on the cheap path. AWS puts the saving at up to thirty per cent. Its boundary is the family: it will not route from one vendor’s small model to another’s large one. It fits where the family you are on already has a cheap member and a capable one. When the split you need crosses families, or depends on business logic outside the prompt, the custom router is the one that fits.
Whichever router runs, measure quality per route. Aggregate accuracy hides the misroutes because they are a minority inside a mostly-correct stream, so sample the requests that took the cheap path and check them against what the capable model produces. That sample tells you whether the threshold is set right, and catches a cheap path whose answers have started to degrade. Then stack the other levers: a cache in front removes repeated requests before any model runs, and trimming the prompt cuts the per-call cost on both paths. Routing, caching, and trimming are separate cuts at the same bill, and they compound.
Worked example
The team samples a day of traffic and sorts it. Roughly sixty per cent is intent classification and short extraction, and twenty per cent is factual questions the knowledge base already answers. The last twenty per cent is the hard tail of multi-condition policy reasoning. Three shapes, and the flagship was serving all of them.
The classification and extraction slice is legible from the endpoint, so the task type alone routes it straight to the small model with no extra call. The factual-question slice goes through the cache and the knowledge base first. Only the residue that needs generation reaches the small model, so most of it never reaches a large-model call at all. The hard policy slice is where difficulty hides inside ordinary-looking questions, so it runs small-model-first with escalation. The cheap model attempts the answer, and a validator checks that every policy condition was addressed. Anything short of that goes to the capable model. The threshold is set to escalate on any unmet condition, because a wrong policy answer reaching a customer costs far more than a spare capable-model call.
After a fortnight the per-route sample tells the story. The cheap path handles the classification and factual slices with accuracy matching the old flagship-only numbers. The policy path escalates about a third of its requests, the tail that genuinely needed reasoning. The flagship now runs on the twenty-odd per cent of traffic that exercises it, and the cache absorbs the repeats. The bill falls by more than half, most of that from moving the easy majority off the big model.
What’s worth remembering
- Model size is a cost lever, not a safety dial, so routing pays off on a mixed workload and does nothing for a uniformly hard or trivial stream.
- Misroutes are asymmetric: overpaying on an easy request loses a few tokens, a hard request on the weak model returns an authoritative wrong answer, so route up when in doubt.
- API-based model cascading adapts to hidden difficulty because the cheap model actually attempts the task, at the cost of hard requests paying for both models.
- Amazon Bedrock Intelligent Prompt Routing pairs exactly two models from one family, keeps traffic on the fallback unless the other clears the response quality difference, and bills USD$1 per 1,000 requests.
- An ensemble raises confidence, not speed: N models multiply the bill by N and finish no sooner than the slowest, so keep it for decisions where being wrong is expensive.
- Measure quality per route, not in aggregate; sample the cheap path against the capable model, because misroutes are a minority hidden inside a mostly-correct stream.