Exam Room · Advanced Generative AI Developer

Picking a Bedrock Model for High-Volume RAG

· 28 min read

Generative AI Development · part of The Exam Room

The situation

A B2B SaaS platform is shipping an in-product assistant. Users ask questions of their own data; the application retrieves relevant records, stitches them into a PromptThe input you hand to an LLM – system instructions, user message, examples, retrieved documents, tool descriptions, the lot., and asks a foundation model to answer. Measured over three months of production traffic:

  • ~1,000,000 requests per day, peaking at 30 RPS during US/EU business-hours overlap.
  • Median request: ~3,000 input TokenThe unit of text an LLM actually sees – usually a short character sequence, not a whole word. (System promptThe instruction block that frames the model’s behaviour for a session, separate from the user’s messages. + retrieved context + user question), ~400 output tokens.
  • P99 first-token latency target < 1.5 s. The UI streams the answer.
  • Quality bar: complex reasoning over structured retrieved context, tables, JSON, pulling answers from multiple documents.
  • Multi-region failover is hard-required. Customers in both us-east-1 and eu-west-1; a regional Bedrock incident must not take either customer base down.
  • Bedrock-native. No separate model-serving infrastructure.

What actually matters

A model choice is a product choice. It fixes who owns the upgrade cadence, who tracks the pricing page, and who gets paged when answer quality drifts after a new model version lands. On a hosted-foundation-model platform those answers split three ways: the vendor ships the behaviour, the platform ships the availability, the team owns the integration. Moving a production RAGA pattern where you retrieve relevant documents at query time and stuff them into the prompt so the model can ground its answer on them. application between model families means rewriting and revalidating the prompt, so the model and the service tier underneath it are one decision, not two.

The second question is what does a bad day look like? At a million requests a day the interesting failure is not an individual bad answer, it is a region going dark for forty-five minutes. The blast radius is every customer homed in that region unless the architecture spreads the load. That pushes the design toward something the application calls with a single model identifier while the platform distributes requests across regions, because the alternative is the application owning a regional routing table and every deploy risking a misrouted call.

Third, the bill. This workload generates roughly 90 billion input tokens and 12 billion output tokens a month, so a difference of one dollar per million input tokens is about USD$90,000 a month on its own. Claude models on Bedrock are billed through AWS Marketplace, and the charges appear under the model provider rather than under Amazon Bedrock, which is worth knowing before anyone goes looking for them in Cost Explorer. Read the current rates off the Bedrock pricing page rather than from memory, then ask which slice of traffic each tier can answer well enough.

Fourth, drift and reversibility. Model behaviour changes between versions, and a system prompt calibrated against one version does not reproduce the same answers on the next. The gap between noticing “answers are slightly worse this week” and measuring “we lost 3% accuracy” is an evaluation pipeline running nightly against a golden set. Geo inference profiles that abstract the specific Region away, prompt templates that separate stable prefix from volatile context, and an evaluation harness that can A/B a new version are what let the team move to a new model in a week rather than six.

What we’ll filter on

Distilling that exploration into filters we can score each model against:

  1. Reasoning quality on retrieved context. Reasoning across long structured prompts, not fluent extraction.
  2. First-token latency under 1.5 s at P99 for ~3,000-token inputs. Tail, not average. AWS publishes no per-model latency figures, so this is a measurement and a service-tier question rather than a spec to read off a page.
  3. Cost per token at ~90B input and ~12B output tokens a month, taken from the current Bedrock pricing page rather than remembered.
  4. Bedrock-native multi-region availability across US and EU, surviving one region offline, without the application routing the calls.
  5. Currency. A model already inside its legacy window is not something to build a two-year product on.

The landscape

Bedrock’s catalogue now spans more than a dozen providers, including Anthropic, Amazon, Meta, Mistral AI, Cohere, AI21 Labs, DeepSeek, Google, OpenAI, Qwen, xAI and Writer. Only a handful clear the reasoning and residency bars together.

Anthropic Claude. Three tiers matter here: Haiku 4.5 (200K context, 64K max output), Sonnet 5 (1M context, 128K max output) and Opus 5 (1M context, 128K max output), with Sonnet 4.6 still active alongside them. All support response streaming, tool calling, image input, prompt caching, guardrails and Bedrock model evaluation. Availability is wide: us, eu and au geo profiles plus a global profile. Note that Sonnet 5 and Opus 5 offer no in-Region option on bedrock-runtime at all, so a geo or global profile is the only way to call them.

Amazon Nova. Nova Micro, Lite and Pro are current; Nova Pro carries a 300K context but only a 5K output cap, and it does have a real EU geo profile (eu.amazon.nova-pro-v1:0) covering Frankfurt, Stockholm, Ireland and Paris. Nova Premier, the tier that would clear the reasoning bar, reaches end of life on 14 September 2026 and has a US geo profile only, so it is not a foundation for new work.

Meta Llama. Llama 3.1 (8B, 70B, 405B), Llama 3.2 (1B, 3B, 11B, 90B), Llama 3.3 70B, and the Llama 4 Maverick and Scout MoE models. Among the lowest pricing on the platform, and a reasonable fit for extraction and summarisation. Llama 3.1 405B and Llama 3.2 90B, the largest of the older generations, both carry July 2026 legacy dates on their model cards, and a model in the Legacy state is closed to new customers. That leaves Llama 3.3 70B and the two Llama 4 models as the selectable top tiers. Llama 4 Maverick has a us. geo profile and nothing else, which is what rules it out here.

Mistral AI. Mistral Large 3 (mistral.mistral-large-3-675b-instruct) is the flagship at USD$0.50 / USD$1.50 per million input / output tokens in US East Regions, with Ministral and Magistral variants below it. Strong multilingual work and solid mid-tier reasoning, below Sonnet on the multi-document reasoning this workload needs.

Cohere. Command R+ was purpose-built for RAG, citation generation and grounded answers, but it reached end of life on 19 August 2026, so it is not selectable. Cohere’s Rerank 3.5 and Embed v4 remain useful in a retrieval pipeline.

Amazon Titan. The family has narrowed to embeddings and image generation. Titan Text Embeddings V2 (amazon.titan-embed-text-v2:0) is active, takes 8K tokens of input, and has configurable output dimensions. It serves the embedding side of a RAG pipeline, not the generation side.

AI21 Labs. Jamba 1.5 Mini and Large, hybrid SSM/Transformer models available in us-east-1 only. Both entered their legacy window in May 2026 with an EOL date of 26 November 2026.

Evaluation

Side by side

Retired and legacy-window models are left out: Command R+ passed EOL in August 2026, Nova Premier and the Jamba 1.5 pair have EOL dates inside the next three months, and Llama 3.1 405B and Llama 3.2 90B are both marked Legacy.

Family Reasoning on retrieved context EU geo profile Cost at 1M req/day Output cap fits Current, not legacy
Anthropic Claude Sonnet 5 ✓ ✓ ✓ ✓ ✓
Amazon Nova Pro ✗ ✓ ✓ ✓ ✓
Meta Llama 4 (Maverick, Scout) ✓ ✗ ✓ ✓ ✓
Mistral Large 3 ✗ ✗ ✓ ✓ ✓

Residency is a weaker filter than it looks, because Nova Pro clears it. What removes Nova Pro is the quality bar: multi-document reasoning over structured context is the tier above it, and that tier is Nova Premier, which is on its way out. Llama’s strongest current models have no EU geo profile. Sonnet 5 is the only row with all five ticks.

Matching the workload to the model

1M req/day RAG 3K input, 400 output, US + EU, P99 < 1.5 s Multi-document reasoning on retrieved context? Nova Pro, Mistral Large 3, Jamba 1.5 (extraction-biased) EU geo profile for EU tenants? Llama 4 (US only); Command R+ past EOL; Llama 3.1/3.2 legacy Measured warm first-token < 1.5 s? Nothing, on paper: AWS publishes no per-model latency. Measure. Cost at 1M req/day tolerable? Opus 5 on rate; easy slice routed to Haiku 4.5 Claude Sonnet 5 us.anthropic.claude-sonnet-5 for US tenants eu.anthropic.claude-sonnet-5 for EU tenants Standard tier + a cache checkpoint over tools and system Haiku 4.5 load-shed fallback; cross-geography retry on 5xx nightly model evaluation against a 500-question golden set
Four gates, reasoning, residency, latency and cost, and the catalogue collapses to Sonnet 5 on a geo profile with Haiku 4.5 as the cost-tier fallback.

The solution

Sonnet 5 is where most production RAG applications land: near-Opus reasoning on realistic prompts, a 1M context window that a 3K input never comes close to, and a 128K output cap. One detail matters for the latency target. Sonnet 5 runs adaptive thinking by default, including on requests that omit the thinking field, and thinking happens before the first visible token. The effort level is configurable and "thinking": {"type": "disabled"} turns it off, so this is a dial to measure against the 1.5 s P99 rather than a fixed cost.

Version choice. Sonnet 5 launched on 30 June 2026, and its model card gives an EOL no sooner than 30 June 2027 with a legacy period of at least six months before that. New applications default to Sonnet 5. A pipeline already calibrated against 4.6 stays there until its evaluation set has been re-run, because a RAG system prompt is tuned against one specific model and the outputs differ between versions.

Service tiers, and the trap in them. Bedrock has four: Standard (the default, pay per token), Priority (a premium for the fastest response times), Flex (a discount for work that tolerates delay) and Reserved (input and output tokens per minute at a fixed price per 1K TPM, 1-month or 3-month terms, billed monthly, overflowing to Standard above the reservation, with minimums of 100,000 input TPM and 10,000 output TPM, arranged through the account team). Tier support is per model, and this is where an architecture drawn from the tier list alone goes wrong: Sonnet 5 supports Standard and nothing else. No Priority, no Flex, no Reserved, and no batch inference either. Haiku 4.5 adds Reserved and batch; Nova Pro adds Priority and Flex. So for this scenario there is no capacity reservation to fall back on and no priority tier to shorten the tail. Peak headroom comes from raised on-demand quotas, and the latency budget has to be met by caching and measurement. Provisioned Throughput still exists as a separate, hourly per-Model-Unit purchase with no-commitment, 1-month or 6-month terms, and it is mandatory for customised models, but it is not the lever for a stock Sonnet 5 deployment.

Prompt caching is the largest single lever on input cost, and it has a minimum most implementations trip over. Cache reads are billed at roughly 10% of the normal input rate and cache writes at roughly 25% above it. Sonnet 5 requires at least 1,024 tokens per cache checkpoint, and allows four per request; below the minimum the request still succeeds and the prefix is not cached, with no error raised. Checkpoints are evaluated in the order tools, then system, then messages, and the minimum applies to the cumulative total across all three, so tool definitions and few-shot examples count toward clearing it. Changing an earlier section invalidates the later ones, so put the checkpoint after the stable block and keep retrieved context behind it. The default TTL is 5 minutes; passing "ttl": "1h" extends it. Haiku 4.5’s minimum is 4,096 tokens, so the fallback tier will not cache a prompt this size at all.

Geo inference profiles turn the multi-region requirement into a config change. us.anthropic.claude-sonnet-5 keeps data within US and Canada Regions, and eu.anthropic.claude-sonnet-5 keeps it within EU Regions, offered from Frankfurt, Zurich, Stockholm, Milan, Spain, Ireland, London and Paris. If one constituent Region fails, the others serve, with no code change. A geography’s destination list never changes once published, which is what makes it safe to name in a compliance answer. The global. profile widens the pool further and drops the residency guarantee, so geo profiles are the right default when US and EU customers are separate.

Cascading for cost. Not every question needs Sonnet 5. Sending short queries and straightforward extraction to Haiku 4.5, the cheaper tier, is where the daily bill bends. Three shapes are common: cascade (try Haiku, re-ask Sonnet when the confidence score is low), pre-route (classify first, choose once), and load-shed (Sonnet by default, drop to Haiku when measured P99 climbs). Cascading degrades most smoothly, because a low-confidence result is re-asked rather than returned. Bedrock’s intelligent prompt routing lists neither Sonnet 5 nor Haiku 4.5 as supported, so the routing is application-side code.

Worked example

  • Primary model: Sonnet 5 via us.anthropic.claude-sonnet-5 for US tenants, eu.anthropic.claude-sonnet-5 for EU. The application routes tenants by home region.
  • Service tier: Standard, with on-demand quotas raised in advance for 30 RPS peak plus margin. No Reserved option exists for this model.
  • Prompt caching: one checkpoint after the tool definitions and system prompt, sized to clear the 1,024-token minimum. Cache hit rate monitored as a first-class metric.
  • Cost-tier fallback: shed to Haiku 4.5 via us.anthropic.claude-haiku-4-5-20251001-v1:0 or the EU equivalent when measured P99 exceeds 2 s for 5 minutes.
  • Cross-geography failover: on repeated 5xx from the primary profile, retry once against the other geography. Degraded-residency mode for continuity.
  • Evaluation: 500-question golden dataset, nightly Bedrock model evaluation run, alert on aggregate drops above 5% week on week.
  • Embedding model: amazon.titan-embed-text-v2:0 in each Region, 8K input limit, dimensions set at invoke time, vector store local to where it is queried.

The monthly volume, which is what the rates get multiplied by:

  • Input: 3,000 x 1M x 30 = 90B tokens. With ~1,200 of the 3,000 input tokens served from cache at roughly a tenth of the input rate, the effective billable input is nearer 58B.
  • Output: 400 x 1M x 30 = 12B tokens, none of it discountable by caching.
  • Routing ~40% of traffic to Haiku 4.5 through a well-calibrated cascade moves that share onto the cheaper rate at similar quality on the easy slice.

Take the current per-million rates off the Bedrock pricing page and multiply. A trustworthy judge is what makes the cascade safe to run, and the Reserved tier counts cache writes as well as ordinary input toward its TPM, which matters if the fallback tier ever gets a reservation.

What’s worth remembering

  1. Check EOL dates first. Command R+ is already retired; Nova Premier and the Jamba 1.5 pair are inside their legacy windows.
  2. Match Claude tier to job. Haiku 4.5 for latency and cost, Sonnet 5 as production default, Opus 5 as escalation.
  3. Geo profiles replace regional routing. us., eu. and au. fail over within the geography; Sonnet 5 and Opus 5 have no in-Region option.
  4. Service tier support is per model. Sonnet 5 is Standard only: no Priority, Flex, Reserved or batch inference.
  5. Cache checkpoints have a minimum size. 1,024 tokens on Sonnet 5, 4,096 on Haiku 4.5; below that nothing caches and no error is raised.
  6. Thinking defaults on in Sonnet 5. It runs even when the thinking field is omitted; set the effort level or disable it.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.