Exam Room · Advanced Generative AI Developer

Multi-Region Resilience for a GenAI Service

· 31 min read

Generative AI Development · part of The Exam Room

The situation

A company runs a customer-facing assistant on Amazon Bedrock in a single Region. The stack is familiar: an application tier behind an API, a Claude model invoked on demand, a guardrail that filters both the prompt and the response, and a knowledge base for retrieval, backed by a vector store and fed from a bucket of source documents. It works well on a normal day.

Two things have started to hurt. At peak, the on-demand model calls hit the account’s per-Region throughput quotas and start returning throttling errors, so users see slow or failed responses when traffic is highest. And the whole assistant lives in one Region, so a recent multi-hour Bedrock disruption there took the feature offline with no fallback. Leadership now wants an availability target the single-Region design can’t meet.

A constraint sits underneath both problems. Some of the traffic carries customer data that, by contract, has to stay within a defined geography. So any answer that spreads load or fails over to other Regions has to respect where those requests are allowed to be processed, not just where they happen to start.

What actually matters

Resilience for a model-backed service is two problems. There’s the throttling problem, which is about capacity and happens on ordinary days, and there’s the regional-failure problem, which is about survival and happens rarely but totally. They have different fixes. Raising effective throughput does nothing for a Region outage, and a warm standby does nothing for a quota you hit at 2pm every Tuesday.

The throttling side is about spreading invocations. On-demand model calls draw on per-Region account quotas, and one Region gives you one set of them. A cross-Region inference profile routes each call to one of several Regions, and it draws on a separate, higher pair of quotas: the cross-Region requests-per-minute and tokens-per-minute values for that model, viewed and raised in the Region you call from. Spikes in one Region are absorbed elsewhere, and the throttling errors thin out.

Two kinds of profile exist, and they differ on exactly the constraint this company has. A geographic profile, prefixed us., eu. or apac., keeps processing inside that geography. A global profile, prefixed global., routes to any supported commercial Region worldwide and prices input and output tokens roughly 10% lower. Either way, prompts and outputs can move outside the Region the request started in, and where AWS stores data for abuse detection it stores it in the destination Region. Routing also has to be permitted. If a service control policy blocks a destination Region the call fails, so the policy needs those Regions allowed, or an inference-profile exception. The destination Regions themselves don’t have to be enabled in the account.

The regional-failure side is about having somewhere to go when a Region is gone. Every Regional piece has to be sitting there already. Model access is a smaller step than it once was: access to Bedrock foundation models is enabled by default in commercial Regions given the AWS Marketplace permissions, and Anthropic’s first-time-use form is one submission per account or organisation rather than per Region, though opt-in Regions need it again. What still varies is whether the model is offered in that Region at all. A guardrail is a Regional resource and is not replicated, so it has to be created and versioned on both sides, and the same goes for prompts and the knowledge base. A failover Region that takes app traffic but can’t invoke the model, or invokes it with no guardrail applied, will not carry the service.

Retrieval is the part people forget. Answer quality depends on the knowledge base, and a knowledge base is Regional. The vector store and the ingestion pipeline both sit in one Region, and an S3 data source has to be in that same Region as the knowledge base. Fail the app and the model over to a second Region holding an empty or stale index, and the service runs while returning worse answers. Making retrieval survive a Region loss means replicating the source documents and keeping a second knowledge base and vector store built and current from them.

Model availability differs by Region. Not every model, and not every feature of a model, is offered everywhere, and new models often land in a subset of Regions first. The failover Region is viable only if it supports the models and features you depend on. That constraint can decide which Region you pair with, and sometimes it decides which model you standardise on.

Then there’s the number that governs the architecture: how much downtime and how much data loss you can tolerate. A tight recovery-time objective, seconds to a minute or two, pushes you toward active-active, because there’s no time to promote anything. A looser one tolerates active-passive, a warm standby you promote when the primary fails. Active-active adds a second live environment and a second operational surface. The standing costs are what double, the app tier, the vector store and the ingestion; token charges follow traffic, so they split rather than double. The recovery objectives, the residency rule, and the cost of a second Region are the three levers.

What we’ll filter on

  1. Availability target: what recovery-time and recovery-point objectives does the service have to hit?
  2. Failure mode covered: peak-hour throttling, full regional outage, or both?
  3. Data residency: which geography or Region are the requests allowed to be processed in?
  4. Component readiness: is the model offered in the second Region, and are the guardrail, prompts and knowledge base present and current there?
  5. Retrieval durability: is the source data replicated and the vector index kept current across Regions?
  6. Cost and operational load: what does keeping the second Region warm, or live, cost in money and in things to run?

The landscape

Single Region, on-demand. The starting point. One Region’s quotas, one Region’s fate. Simplest to run, cheapest, and the residency story is trivial because nothing leaves. It can’t meet a demanding availability target and it caps out at one Region’s throughput.

Provisioned Throughput. Purchase model units for one model in one Region, billed hourly, with no commitment, a one-month term or a six-month term. It removes on-demand throttling for a sustained load and gives consistent latency. Three things rule it out here. It’s a Regional purchase and does nothing for a Region outage. Inference profiles don’t support Provisioned Throughput, so it can’t be combined with cross-Region routing. And the published list of base models available for purchase stops at earlier generations, Claude 3.5 Sonnet v2 being the newest Anthropic entry, so a current Claude assistant can’t buy capacity for the model it actually runs. It remains the route for custom models, which can only be served this way.

Geographic cross-Region inference profiles. Invoke through a us., eu. or apac. profile instead of a single model ID, and Bedrock routes each call to one of that geography’s Regions against a separate, higher quota. This answers the peak-hour capacity problem and needs no standby infrastructure of your own. The residency consequence is that a request can be processed in another Region inside the geography, so the geography has to sit within the area the data is allowed to be in.

Global cross-Region inference profiles. The same mechanism with the boundary removed: a global. profile routes to any supported commercial Region worldwide, and input and output tokens price roughly 10% below the geographic equivalent. Support is narrower, currently Claude Sonnet 4.5 from a set of source Regions. The worldwide routing is what disqualifies it here, and an organisation that needs it disabled denies the request condition aws:RequestedRegion of unspecified.

Active-passive (warm standby) across Regions. A second Region kept ready with the guardrail, the prompts and a current knowledge base, but not taking live traffic until the primary fails, at which point routing shifts. This survives a full regional outage at a moderate cost, with a recovery time measured in the minutes it takes to detect and cut over. It needs the standby kept warm, and the replication running continuously so the index isn’t stale when you land on it.

Active-active across Regions. Both Regions serve live traffic all the time, fronted by health-checked routing, so a failure stops traffic reaching the sick Region. This meets the tightest recovery objectives because there’s nothing to promote, and it doubles as capacity headroom. The cost is two standing stacks and two live environments to keep in step: guardrails, prompts and knowledge bases identical on both sides, which is real operational work.

Front-door routing (Route 53 / Global Accelerator). Not a standalone answer but the mechanism that makes failover work. Route 53 health checks run at a 30-second standard interval or a 10-second fast interval, and flip state after a configurable number of consecutive failures, so detection takes tens of seconds before the record TTL is added on top. Global Accelerator instead publishes static anycast IP addresses and reroutes behind them, so there’s no DNS cache to wait out.

Evaluation

Side by side

Option Fixes throttling Survives Region outage Recovery time Residency control Steady-state cost
Single Region, on-demand ✗ ✗ N/A (down) Full (nothing leaves) Lowest
Provisioned Throughput ✓ (listed models only) ✗ N/A (down) Full (one Region) Hourly per model unit
Geographic inference profile ✓ ✗ (invocation only) N/A Geography-bounded Usage-based
Global inference profile ✓ ✗ (invocation only) N/A ✗ (routes worldwide) Usage-based, ~10% lower
Active-passive (warm standby) Partial ✓ Minutes (promote) You choose Regions Medium
Active-active ✓ ✓ Health checks plus TTL You choose Regions Two standing stacks
Front-door routing ✗ Enables it Drives the cutover Neutral Low

The table splits the two problems cleanly. A cross-Region inference profile answers throttling but not outage, because it spreads the model invocation and nothing else. Active-passive and active-active answer outage, and which one fits depends on the recovery-time objective and the budget. They compose: a real design runs a geographic profile for throughput inside whichever multi-Region posture the availability target demands.

Choosing a multi-Region resilience posture for a Bedrock service Four rows read left to right. Each of four needs on the left connects to a gate in the middle and an answer on the right. The first three rows lead to a cross-Region inference profile, active-active Regions, and an active-passive warm standby. The fourth row is the data-residency rule, which bounds the Region choice for the other three. What you need Higher throughput, throttling smoothed at peak Survive a full Region outage, tight recovery time Survive a Region outage, minutes of recovery is fine Data must stay in a defined geography (applies to every option) Gate Invocation-only spread, no standby to run? No time to promote a standby on failure? Warm standby, promote on failover? Geographic profile, not a global one Pick Cross-Region inference profile Active-active Regions, health-checked routing Active-passive warm standby, failover routing Bounds Region choice for all three above yes yes yes

The solution

Fix throttling first, because it’s the cheaper problem and needs no second environment of your own. Switch the model invocation from a single model ID to a geographic cross-Region inference profile whose geography covers the Regions you’re entitled to use. Each call then routes across that geography’s Regions against the higher cross-Region quotas, and the peak-hour throttling errors clear without provisioning anything. Check the service control policies before you cut over, because a policy that blocks a destination Region will fail the call. Two caveats decide whether the profile is safe. Traffic that is contractually pinned to one Region can’t ride a profile that would route it out, so that class stays on a direct in-region invocation. And a profile addresses the model call only, so it does nothing for a Region outage while the app tier, guardrail and knowledge base live in one place.

For surviving a Region outage, the recovery-time objective picks the posture. If the target is minutes and the budget is moderate, build active-passive. The second Region needs the model offered there, the guardrail recreated and versioned with the same configuration, the prompts present, and a standing knowledge base with its own vector store. Route 53 health checks watch the primary, and a failover routing policy shifts users across when it goes unhealthy. Keep the record TTL short, because the TTL is added to the detection time. What makes this real is keeping the standby current continuously, not on the morning of the incident.

If the target is tighter than a promote-and-cut-over can hit, go active-active. Both Regions serve live traffic behind latency or weighted routing, and a failed Region stops receiving requests. There’s nothing to promote, so recovery is the time for health checks to react plus the TTL, and Global Accelerator removes the TTL term by putting static anycast addresses in front. The costs are two standing stacks and the work of keeping two live environments identical: the same guardrail configuration, the same prompt versions, and knowledge bases that answer the same way. Drift is the failure mode, so the guardrail, prompt and ingestion config should deploy from one source of truth rather than being clicked into place twice.

Whichever outage posture you choose, retrieval has to come along, or the failover returns worse answers. A knowledge base can only read an S3 data source in its own Region, so replicate the source bucket with S3 Cross-Region Replication and point a second knowledge base and vector store at the replica. Newly added documents replicate and then get ingested on the other side, which keeps the recovery-point gap to the replication and ingestion lag rather than a full rebuild. Standard replication is asynchronous with no time guarantee; S3 Replication Time Control replicates 99.9% of objects within 15 minutes and publishes metrics and threshold events, which turns that gap into a number you can report. Confirm, before committing to a Region pairing, that the second Region offers the models and features you depend on.

Worked example

The assistant starts in a single Region: app tier, on-demand Claude invocation, one guardrail, one knowledge base over an OpenSearch Serverless vector store fed from an S3 bucket. Its contract says customer data must stay within one geography. The new availability target allows a couple of minutes of recovery, and peak traffic throttles the model calls.

The throttling fix lands first. The app stops calling a single model ID and calls the geographic inference profile for that geography, so invocations spread across its Regions against the cross-Region quotas and the peak-hour throttling clears. The global profile is rejected here despite the cheaper tokens, because it routes worldwide. The service control policy is updated to allow the profile’s destination Regions.

The outage fix is active-passive, because minutes of recovery is acceptable and it costs less. A second Region in the same geography is confirmed to offer the model, then gets the guardrail recreated with the identical configuration and version, the prompt library deployed, and a knowledge base with its own vector store. S3 Cross-Region Replication with Replication Time Control copies the source documents across, and the second knowledge base ingests them so its index stays current. Route 53 health-checks the primary endpoint on the fast interval, with a short TTL on a failover record pointing at the standby. The recovery-point gap is the replication and ingestion lag; the recovery time is detection, TTL, and the routing switch. When the primary Region has its next bad hour, traffic moves to a standby that has the model, the guardrail, the prompts and a warm index, and the assistant keeps answering from inside the geography it’s allowed to run in.

What’s worth remembering

  1. Throttling and regional outage are two problems: spreading invocations fixes capacity, a second Region fixes survival, and neither fix touches the other.
  2. A cross-Region inference profile draws on separate, higher cross-Region quotas for the model, managed in the Region you call from, and needs no standby of your own.
  3. Geographic profiles (us., eu., apac.) keep processing inside a geography; a global. profile routes worldwide for roughly 10% less per token, which residency rules usually rule out.
  4. Provisioned Throughput is a Regional purchase, can’t be used with an inference profile, and covers only the models on the published list, so it isn’t the throughput answer for a current model.
  5. Guardrails, prompts and knowledge bases are Regional and aren’t replicated, and a knowledge base reads an S3 data source only in its own Region, so failover needs them rebuilt and kept current.
  6. Route 53 failover costs detection time plus the record TTL; Global Accelerator’s static anycast addresses remove the TTL term.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.