The situation
A product team runs a customer-facing assistant on Amazon Bedrock, calling one Claude model on-demand. It was comfortable in testing and through the first month of light traffic. Then a marketing push doubled sign-ups, an overnight batch job that summarises the day’s tickets started overlapping with daytime interactive load, and the logs filled with ThrottlingException. Some requests now fail outright; others succeed only after several seconds of silent retrying, and the p99 latency has crept from under two seconds to well over ten.
The team’s first instinct was a tight retry loop that re-sends the call until one succeeds. That made the failures quieter but the latency worse, because every throttled request now spends its time re-queuing rather than erroring fast. Nobody has looked at whether the account is actually over its Bedrock quota, or whether the batch job and the interactive traffic even need to share the same capacity at the same moment.
Bedrock meters on-demand inference in tokens. Each model has a tokens-per-minute and a tokens-per-day quota in each Region, and the deduction happens when the request arrives: input tokens plus whatever max_tokens the call asked for, with the unused remainder returned once the response completes. Some models also carry a requests-per-minute quota. Several recent Claude models carry none, and are governed by tokens alone. Cross a ceiling and the service returns ThrottlingException. Underneath the noise sits one distinction: are these transient bursts a good client can ride out, or a structural shortfall no amount of retrying will fix?
What actually matters
Throttling is a symptom, and the same symptom has three different causes that need three different fixes. Reaching for retries first is right for one of them and useless for the other two, so the first thing to establish is which one you’re looking at.
A transient throttle is a short burst that briefly exceeds the per-minute ceiling while average demand sits comfortably under it. This is what client-side resilience is for. Exponential backoff with jitter spreads the retries out so the burst drains and the retried requests land in a quieter window; the average was always servable, the arrivals were just clumpy. The AWS SDKs do this for you. Standard mode, the default, retries throttling and transient errors with exponential backoff and full jitter, waiting from a one-second base on a throttle and capping any single wait at twenty seconds. It also holds a retry token budget, so when a large share of calls is failing the client returns errors immediately rather than queueing behind attempts unlikely to succeed. Retries turn a jittery arrival pattern into a smooth one, and add little latency while the underlying capacity is adequate.
A structural throttle is different. Sustained demand genuinely exceeds the quota, and no retry strategy adds a single token per minute of capacity. Retrying a structurally throttled workload converts fast failures into slow ones and, if the whole fleet backs off and retries in step, into correlated stampedes. The fixes change the capacity, not the client. You can request an increase to the model’s token quotas in that Region through Service Quotas. You can move the traffic onto a cross-Region inference profile, where Bedrock selects a Region to process each request and the calls draw on a separate cross-Region token quota instead of the single-Region on-demand one. Or you can reserve capacity through Bedrock’s Reserved service tier, which holds a stated input and output tokens-per-minute floor and overflows to the Standard tier above it. Each of these raises the ceiling. Backoff never does.
Then there’s the shape of the demand itself, which you can change without touching the ceiling. A lot of throttling is self-inflicted synchronisation: a batch job that fires a thousand requests at once, or interactive and background work colliding in the same minute. Putting non-interactive work behind an SQS queue drained at a controlled concurrency turns a spike into a steady stream that fits under the quota. It also separates the batch job’s timing from the interactive path, so the two stop drawing on the same per-minute budget in the same minute.
Two things decide which lever fits: latency tolerance and criticality. Interactive requests have seconds of budget at most, so their answer to overload is fast backoff then graceful degradation, shedding or deferring the request rather than making a user wait a minute. Background work has a generous latency budget, so it can absorb queueing and long backoffs invisibly. Bedrock’s service tiers encode that split directly: a service_tier of priority puts a request ahead of standard and flex traffic, flex takes a discount in exchange for longer processing, and all three draw on the same on-demand quota. And when capacity is genuinely scarce, criticality decides what gives. Shed or defer the low-priority work, and optionally fall back to a smaller model with separate quota and a lower price, keeping the important path answered while the nice-to-have path waits.
What we’ll filter on
- Transient or structural? Is the average demand under the quota with clumpy arrivals, or genuinely over the ceiling?
- Latency tolerance, does this request have seconds to answer or minutes?
- Criticality, is this interactive work that must be served, or deferrable background work?
- Adds capacity or just reshapes arrivals? Does the lever raise the ceiling, or smooth the traffic under it?
- Time and commitment to apply, an SDK setting today versus a quota request or a monthly reservation.
The landscape
Exponential backoff with jitter (SDK retries). The first line for transient throttles. On a ThrottlingException the SDK waits a growing, randomised interval and retries, so retries from many callers don’t all fire at the same instant. Standard mode is the default, and the right mode for a latency-sensitive caller. Adaptive mode adds a client-side rate limiter that can delay the initial request as well as the retries; AWS recommends it for a client calling a single resource hard and tolerant of latency, and advises against it as a general default. The 2026 backoff timings and retry-quota behaviour are opt-in until they become the default, through AWS_NEW_RETRIES_2026=true. Immediate to apply, and it adds no capacity: point it at a structurally over-quota workload and it only slows everything down.
Circuit breaker. Backoff assumes the dependency recovers in a few hundred milliseconds. A breaker is for when it does not, and it stops the caller waiting out a timeout on every request while the dependency is down. The pattern gives the caller three states. Closed, where calls pass through and failures are counted against a threshold. Open, where the breaker has tripped and calls fail fast into the degraded path without reaching Bedrock at all. Half-open, where a single probe request determines whether to close it again. Where the state lives determines whether the pattern works at fleet scale. A per-container in-memory counter never sees the fleet’s failure rate and resets on every cold start. Put the flag in a DynamoDB item or an AWS AppConfig value instead, keyed per model and per Region, read by every Lambda that calls the model. In Step Functions the breaker becomes an explicit state rather than a library: a Choice on the stored health flag ahead of the model task, a Retry block on the task itself with MaxAttempts, BackoffRate and JitterStrategy set to FULL (it defaults to NONE), and a Catch that routes to the fallback branch.
A client-side concurrency ceiling. Managing concurrent invocations starts from a number you can calculate. Take the model’s tokens-per-minute quota, divide by the tokens one call reserves (its input tokens plus its max_tokens), and multiply by the mean call duration in minutes. That gives the number of in-flight calls the quota can actually serve. A bounded worker pool or a semaphore in front of the Bedrock client holds the application to that number, so back-pressure lands in your own queue where you can see it, instead of arriving as ThrottlingException retries that consume quota and add latency without adding capacity. Retries and a cap solve different halves of the problem. Backoff handles the transient collision; the cap stops the workload generating collisions at all. Adaptive retry mode without a cap only moves the queue into the SDK. The same sum shows where trimming tokens helps. An oversized max_tokens reserves quota the reply never uses, an output token counts against the quota at a per-model burndown rate, five, ten or fifteen to one on the Anthropic models, and tokens read from a prompt cache don’t count against it at all. Upstream of the client entirely, API Gateway can rate-limit before a request reaches Bedrock: a usage plan gives each API key its own rate and quota, and per-method throttling caps the rate any one route can sustain.
Service tiers. A service_tier parameter on the runtime call selects Reserved, Priority, Standard or Flex. Priority is served ahead of standard and flex requests, at a premium over standard on-demand pricing. Flex takes a discount in exchange for longer processing, which suits evaluations, summarisation and anything with nobody waiting on it. Priority, Standard and Flex all draw on the same on-demand quota, so tiering changes who gets served first under contention rather than adding capacity.
Service-quota increase. Raise the model’s token quotas for a Region through Service Quotas. The direct fix when demand has outgrown the default and you want more of the on-demand pool. It’s a request, not a switch, so it takes lead time and isn’t guaranteed: AWS gives priority to accounts already consuming their existing allocation, and declines increases for models in a Legacy or Deprecated lifecycle status. The cross-Region tokens-per-minute, on-demand tokens-per-minute and tokens-per-day quotas are raised together off one request. It still leaves you on shared on-demand capacity with no reserved floor.
Cross-Region inference profile. A profile that lets Bedrock choose the Region that processes each request, either within a geography such as US or EU, or anywhere in the commercial Regions with a global profile, the latter at roughly ten percent below standard pricing. Calls against a profile draw on a separate cross-Region token quota, and the routing spreads load across more compute than one Region holds. Traffic stays on the AWS network either way, but only a geographic profile keeps processing inside a boundary, so residency rules decide which kind you can use. Inference profiles don’t support Provisioned Throughput.
Reserved capacity. The Reserved tier holds prioritised capacity for a model, with input and output tokens per minute reserved separately so the reservation matches the workload’s shape. Traffic above the reservation overflows to the Standard tier rather than failing. It targets 99.5% uptime for model response, reserves for one or three months at a fixed monthly price per 1,000 tokens per minute, and starts at 100,000 input and 10,000 output tokens per minute, arranged through your AWS account team. The fit is steady, high-volume, latency-sensitive traffic, not bursty or experimental load. For a custom model, or one of the older base models still on the list, the equivalent mechanism is Provisioned Throughput, billed hourly against model units.
Queue and controlled concurrency (SQS). Put non-interactive work behind a queue and drain it with a bounded number of workers, so a thousand-at-once batch becomes a steady stream that fits under the quota. Smooths demand and separates background timing from the interactive path. Bedrock’s own batch inference goes further for bulk work: prompts go to S3 as a job, results come back to S3, and the job runs against quotas separate from the per-minute on-demand ones. Batch jobs don’t support tool calling or structured output, and both mechanisms add latency by design, so they suit deferrable work rather than a user waiting on a reply.
Graceful degradation and model fallback. When capacity is genuinely scarce, shed or defer low-priority requests, and optionally fall back to a smaller or alternate model with separate quota and a lower price. Keeps the important path answered under load instead of failing everything equally. The fallback model needs to be good enough for the degraded path, and you need a clear rule for what counts as low priority.
Evaluation
Side by side
| Lever | Fixes transient | Fixes structural | Adds capacity | Reshapes arrivals | Latency added | Time to apply |
|---|---|---|---|---|---|---|
| Backoff + jitter (SDK) | ✓ | ✗ | ✗ | ✓ | Seconds (on retry) | Immediate |
| Circuit breaker | Partly | ✗ | ✗ | ✓ | Saved, not added | Hours to build |
| Client-side concurrency cap | ✓ | Partly | ✗ | ✓ | Held in your queue | Hours to build |
| Priority / Flex service tier | Partly | ✗ | ✗ | ✓ | Lower, or higher | One parameter |
| Service-quota increase | ✓ | ✓ | ✓ | ✗ | None | Days (request) |
| Cross-Region profile | ✓ | ✓ | ✓ | ✓ | Negligible | Hours to set up |
| Reserved tier | ✓ | ✓ | ✓ | ✗ | None | 1 or 3 months |
| Queue or batch inference | ✓ | Partly | ✗ | ✓ | Minutes to hours | Hours to build |
| Degradation / fallback | ✓ | ✓ | ✗ | ✗ | None (sheds instead) | Hours to build |
Reading the table against the situation: the interactive path needs backoff plus, if the average is genuinely over quota, a quota increase or a cross-Region profile, with degradation as the safety valve. The overnight batch job belongs behind a queue so it stops colliding with daytime traffic. And if the interactive baseline is both high and steady, the Reserved tier gives it a floor of guaranteed tokens per minute, with traffic above the reservation overflowing to standard on-demand. No single lever covers all of it.
The solution
Start by measuring, because the transient-versus-structural split decides everything and the error count alone won’t tell you which one you have. InvocationThrottles in the AWS/Bedrock namespace counts what was rejected. InputTokenCount, CacheWriteInputTokenCount and OutputTokenCount say what you consumed, with the output multiplier applied. There’s an EstimatedTPMQuotaUsage metric as well, and AWS documents it as an approximation rather than a basis for capacity planning, because throttling keys on the reservation made at the start of the request rather than the final count. If average consumption sits under the ceiling and the throttles cluster in short spikes, it’s transient, and backoff is the whole answer. If the average bumps the ceiling for sustained stretches, it’s structural, and no client change will help.
For the transient case, use the SDK’s retry support rather than writing your own loop. Standard mode gives you exponential backoff with full jitter, plus a retry budget that stops a failing fleet retrying itself into the ground. Keep max attempts low and bound the total wait, so an interactive request fails fast enough to degrade rather than hanging; unbounded retries are how a throttle becomes a latency incident. Adaptive mode can delay the initial request, so it belongs on the batch caller rather than the path a user is waiting on. Two small habits belong next to it. Size the client’s concurrency ceiling from the quota rather than letting the worker count drift up with the instance size. And hold one long-lived Bedrock client with a warm connection pool for the life of the process instead of constructing one per invocation, which saves a TLS handshake on every call.
For the structural case, pick the capacity lever by traffic shape. Bursty or still-growing traffic needs a quota increase and, where residency allows, a cross-Region inference profile to spread the load. Both raise the effective ceiling without a monthly commitment. Steady, high-volume, latency-sensitive traffic that can’t tolerate on-demand throttling at all is the case for a Reserved tier reservation, sized from measured input and output token rates, with prompt-cache writes counted in the input figure. The trap is reserving for spiky or experimental load and then holding a one- or three-month minimum that sits mostly idle. For a custom model the same reasoning runs through Provisioned Throughput, where the model units are sized from measured demand rather than guessed.
The batch job is a demand-shape problem, not a capacity one. It fails because it fires everything at once and collides with interactive traffic. So the fix is a queue draining at controlled concurrency, which flattens the spike into a stream that fits under the quota and stops the two workloads competing for the same per-minute budget. It adds latency, which an overnight summarisation job absorbs and an interactive request cannot, which is why the two belong on different mechanisms. Further along the same line, work with no reader until the morning goes to Bedrock batch inference or an asynchronous, event-driven pipeline, drained at whatever rate the quota leaves spare.
Degradation is the safety valve underneath all of it, because even with the right capacity lever a big enough spike can still exceed the ceiling. It works when the ladder is written down before the incident rather than improvised during it, in the order the application will walk it: the full answer from the primary model, then a smaller model with its own separate quota, then a cached answer to the same question from earlier, then a deterministic template reply assembled from what the application already knows without calling a model, then queue the request and answer it later, then an honest error. Each rung is faster and less capable than the one above, and the application only steps down when the rung above has failed or its breaker is open. The degraded reply is still a reply. It is a shorter answer, or the account facts the assistant can state without inference, with a line saying the service was busy. Two things break fallback in practice. On a cross-Region inference profile, track the breaker per Region: one unhealthy Region that opens a single shared breaker sheds traffic the healthy Regions would have served. And a breaker wrapped around an unbounded retry loop is the same as no breaker, since the call never fails quickly enough to trip it.
Worked example
The overnight job summarises the day’s tickets. It reads a few thousand rows and, in a tight loop, fires a Bedrock request per ticket as fast as the code can iterate. Most nights it finishes before the interactive traffic wakes up. On the night the marketing emails go out at 2am local time, early-riser users start chatting to the assistant while the batch is still running, and both streams hit the same model’s per-minute token quota at once. Interactive requests throttle, the tight retry loop on the interactive path spins, and users watch a spinner for fifteen seconds.
The measurement shows the daytime interactive average sits comfortably under quota. The problem is purely that the batch spikes into the same minute. So the batch job goes behind an SQS queue drained by a small, fixed pool of workers, sized so its steady token rate leaves headroom under the ceiling for interactive traffic, and its calls are marked flex so anything contended is served after the interactive ones. The spike becomes a stream, the batch finishes an hour later than before (nobody notices; it’s a summary that’s read at 9am), and interactive throttling on collision nights disappears.
On the interactive path itself, the hand-rolled retry loop is replaced with the SDK’s standard retry mode at two attempts, capped so a request that can’t be served in a couple of seconds fails over to a degraded reply (“we’re busy, here’s a shorter answer”) backed by a smaller model with its own quota. Two fixes for two different causes: the queue reshapes the demand that was colliding, the backoff and degradation handle the residual bursts on the path that can’t wait. Neither of them is a bigger retry loop, and neither would have worked in the other’s place.
What’s worth remembering
- Throttling has three causes, transient bursts, structural over-quota demand, and colliding traffic shapes, and each calls for a different fix; identify which before reaching for a lever.
- Exponential backoff with jitter is the right first response to a transient throttle, and it adds no capacity, so it does nothing for a workload that’s genuinely over quota.
- Cap retries and total wait on interactive paths, so a throttle fails fast into degradation instead of turning into a latency incident.
- Bedrock’s quotas are token quotas, deducted as input tokens plus
max_tokenswhen the request arrives, so an oversizedmax_tokensthrottles you earlier than the replies warrant. - Structural shortfalls are fixed by raising the ceiling: a quota increase, a cross-Region inference profile with its own separate quota, or a Reserved tier reservation for a guaranteed floor.
- A queue, a batch job or the Flex tier reshapes demand rather than adding capacity, and decouples background work from the interactive path.