The situation
A knowledge-assistant team has shipped a Bedrock-backed feature that answers staff questions over an internal document set. A request retrieves passages from a vector store, stuffs them into a prompt with the conversation history, calls the model, and sometimes makes a tool call to look up a live figure before answering. It works, and users say it feels sluggish. The complaint is vague: sometimes it is fine, sometimes it hangs for what feels like an age before anything appears.
The team has one number to go on, an average end-to-end time of about 4.2 seconds, and they have been arguing about it from taste. One camp argues for a bigger vector index with more results per query, sure the retrieval is thin. Another argues for a larger, smarter model, sure the answers are the bottleneck. A third has been told streaming will fix everything and pushes to ship it first. Nobody has measured where the 4.2 seconds goes. The average hides what users are reacting to, which is the occasional request that takes twelve seconds while the rest take two.
The real task is not choosing a lever. It is finding out which stage owns the time, at the tail as well as the middle, and then choosing the lever that stage responds to.
What actually matters
An end-to-end generative-AI request is a pipeline of stages that each add wall-clock time, and they do not add up the way people assume. The stages are retrieval (embedding the query, searching the index, fetching passages), prompt assembly, the model call itself, any tool calls the model triggers mid-generation, and the network hops between them. The model call is not one number either. AWS splits it into prefill, where the model processes the whole input in a single forward pass and produces the first output token, and decode, where every further token takes its own forward pass. The two respond to completely different levers.
Prefill shows up as time-to-first-token, and it scales primarily with input length, rising further when the endpoint is contended. A long prompt, a big pile of retrieved context, and a large model all push it up. Decode time is set by model size and by how many tokens the model produces, so a verbose answer takes proportionally longer whether or not the user reads it all. Per-token decode time is fairly stable for a given model and host load, which makes a drop in output tokens per second a clean signal of service-side slowdown. Hold that split in your head. It explains why streaming improves how the feature feels without touching total time, and why cutting output length cuts total time without moving the first-token wait much.
The tail is what users feel, and an average erases it. If p50 is two seconds and p99 is twelve, the average reads a comfortable-looking four, and the twelve-second requests are the ones generating complaints and abandoned sessions. Measure latency per stage as percentiles, at least p50 and p99. That shows the typical path and the bad one together, and it tells you whether the tail lives in retrieval, in the model, in a tool call that occasionally times out, or in a cold dependency. Attacking the mean optimises the wrong request.
Under all of it sits latency against quality. The fastest single change is almost always a smaller model, and a smaller model can answer worse. The goal is the lowest latency that still clears the quality bar the feature needs. Every latency win from swapping models has to be checked against an evaluation of the output, not just the stopwatch.
What we’ll filter on
- Which stage owns the time, retrieval, prompt assembly, first-token wait, generation, tool calls, or network?
- p50 versus p99 per stage, is the pain in the typical path or the tail?
- Perceived versus total latency, does the user need the answer faster, or just to see it start sooner?
- Input size versus output size, is the cost in what the model reads or in what it writes?
- Quality headroom, how much answer quality can this feature give up for speed before it fails its job?
- Contention and reuse, is the endpoint queuing under load, and does a shared prompt prefix repeat across calls?
The landscape
The levers each target a specific stage, and naming the stage each one touches is most of the skill.
Streaming cuts perceived latency, not total latency. Instead of waiting for the full response, you stream tokens as they generate, so the user sees the first words at time-to-first-token rather than after the whole answer is written. On Bedrock that is ConverseStream or InvokeModelWithResponseStream. Total generation time is unchanged, but a two-second wait that starts producing text at 400 milliseconds feels far faster. Streaming also changes what you can measure: Bedrock publishes the TimeToFirstToken metric only for those two streaming operations, so a non-streaming call gives you a single end-to-end number and no split. For a chat-shaped feature this does more than any other single change, and it does nothing measurable for a batch job nobody watches.
A smaller or distilled model cuts both prefill and decode time, because a smaller model reads and writes faster. This is the biggest single lever on raw latency and the one with the sharpest trade-off, since the smaller model may answer worse. Model families on Bedrock span this range deliberately: a fast, small model for the latency-sensitive path and a larger one where answer quality justifies the wait. The usable version of this lever is routing, sending easy requests to the small model and only the hard ones to the large one.
Fewer output tokens cuts total time directly, because decode is linear in tokens produced. Tighten the prompt to ask for a shorter answer, cap max_tokens, and stop the model restating the question. Capping max_tokens does a second job that is easy to miss. Bedrock deducts input tokens plus max_tokens from your tokens-per-minute quota when the request starts and replenishes the unused part at the end, so a max_tokens of 32,000 on a 1,000-token answer reserves quota nothing uses and throttles you earlier. On several Claude models an output token also burns down quota at a multiple of an input token. Output length therefore sets the throttling ceiling as well as the clock.
Prompt caching cuts prefill time when a large chunk of the prompt is identical across calls. Bedrock prompt caching comes in two forms. Implicit caching reuses an eligible prefix with no change to the request; explicit caching marks the prefix with a cachePoint, so a long system prompt, a fixed instruction block, or a document reused across a session is not reprocessed on later calls. Minimums vary by model, from 512 to 4,096 tokens per checkpoint, with up to four checkpoints per request on Claude models. The cached entry has a time to live that resets on every hit, five minutes by default and an hour on several models. Cache reads are billed at a reduced rate and do not count against the tokens-per-minute quota. The saving lands on input processing, so it helps most when the stable prefix is large next to the variable part, it does nothing for output generation, and batch inference does not support it.
Pre-computation takes the model out of the request altogether for the head of the query distribution. Most features have a small set of questions that account for a large share of traffic: the same dozen policy lookups, the same few phrasings of what the current allowance is. Generate those answers offline on a schedule and serve them from DynamoDB or ElastiCache, and a matching request becomes a key lookup in single-digit milliseconds that never reaches the model. The trade is that a pre-computed answer is stale by construction, so it needs an invalidation trigger tied to the source data changing rather than a time-to-live chosen by feel. It is worth building only where the query set is genuinely predictable, which is what separates it from a response cache that fills opportunistically from whatever traffic happens to arrive.
Trimming retrieved context and prompt size cuts prefill time by giving the model less to read. Retrieval that returns twenty passages when three would do inflates the input, and every extra token is time before the first output token and money on the bill. RerankingA second pass that re-scores a wide set of retrieved candidates and keeps only the few most relevant, so the expensive model reads less. to the few passages that actually matter, and cutting conversation history to what the turn needs, shrinks the input the model must process. Smaller prompts are faster prompts, and cutting irrelevant passages often improves the answer as well.
Running retrieval and other work in parallel cuts total time when stages are independent. If a request needs a vector search and a separate metadata lookup, and neither depends on the other, running them concurrently makes the pair take the time of the slower one rather than the sum. The same applies to independent tool calls, and to workflows where several model calls each produce part of one answer and are joined at the end. Anything on the critical path that does not depend on an earlier result is a candidate to move off the serial chain.
Latency-optimised inference serves the request on infrastructure tuned for speed, cutting prefill and decode time on the same model rather than dropping to a smaller one. You set performanceConfig.latency to optimized on the runtime call; the default is standard. Treat it as a narrow option. AWS still documents it as a preview feature, it covers a short list of models in a few US Regions, it is reached through cross-Region inference, and it costs more per token. Once you exhaust the latency-optimisation quota for a model, Bedrock serves the request at standard latency and charges standard rates, so it is not a capacity guarantee. Reach for it when you have hit the quality floor and cannot shrink the model further.
Cross-Region inference profiles address the tail that comes from contention. An inference profile distributes invocations across the Regions it defines, either inside a geography such as US, EU or APAC, or globally, and cross-Region calls draw on separate, larger requests-per-minute and tokens-per-minute quotas than a single-Region call. That is throttling relief rather than a latency feature. It does little for one uncontended request, and it flattens the p99 spikes that come from saturation at peak, which is often where the twelve-second tail lives. Inference profiles do not support Provisioned Throughput, so the next lever is an alternative to this one, not an addition.
Provisioned Throughput reserves model capacity so requests stop drawing on the shared on-demand pool. You purchase model units, each delivering a set number of input and output tokens per minute, billed hourly with no commitment, a one-month term, or a six-month term. It is a capacity and cost decision more than a per-request tweak, and it suits steady, high-volume, latency-sensitive traffic rather than spiky low-volume workloads. Serving a customised model requires it regardless.
Evaluation
Side by side
| Lever | Stage it targets | Total latency | Perceived latency | Tail (p99) | Quality risk |
|---|---|---|---|---|---|
| Streaming | Prefill to display | ✗ | ✓ | ✗ | none |
| Smaller / distilled model | Prefill and decode | ✓ | ✓ | ✓ | high |
| Fewer output tokens | Decode, and quota reservation | ✓ | ✓ | ✓ | some |
| Prompt caching | Prefill (input reuse) | ✓ | ✓ | ✗ | none |
| Pre-computation | Whole request (predictable queries) | ✓ | ✓ | ✓ | staleness |
| Trim retrieved context | Prefill (input size) | ✓ | ✓ | ✗ | can improve |
| Parallel retrieval / tools | Independent stages | ✓ | ✓ | ✓ | none |
| Latency-optimised inference (preview) | Prefill and decode | ✓ | ✓ | ✓ | none |
| Cross-Region inference profile | Contention | ✗ | ✗ | ✓ | none |
| Provisioned Throughput | Queuing under load | ✗ | ✗ | ✓ | none |
The table reads as a diagnosis tool. If answers start too late, the streaming and prefill levers own it. If the tail is slow under load, the contention levers own it. If the raw number is too high everywhere, the model and input-size levers do the work.
Before any of that, a picture of where the seconds actually go on this feature’s slow path:
The solution
Start by instrumenting, because everything else is guessing until the stages are measured. Wrap each stage in timing, emit the durations, and aggregate them as percentiles rather than averages. CloudWatch holds the per-stage p50 and p99, and the AWS/Bedrock namespace supplies the model-side numbers: InvocationLatency for the whole call, TimeToFirstToken for the prefill wait, InputTokenCount and OutputTokenCount for the sizes, and InvocationThrottles for the requests that never ran. One catch shapes the order of work. TimeToFirstToken is published only for the streaming operations, so a team on non-streaming Converse has to move to ConverseStream before the split is visible at all. Model invocation logging is a separate feature and records request and response bodies with their token counts, not stage timings. The output of this step is a bar per stage at p50 and p99, and it usually settles the argument the team was having by taste. In the picture above, retrieval is not the problem the retrieval camp thought it was; the tail lives in the model call and an occasional slow tool.
Instrumentation gives the diagnosis; a benchmark set turns each lever into a number. Assemble a fixed set of representative prompts drawn from real traffic, covering the short factual questions as well as the long multi-passage ones. Replay it at a target concurrency before and after every change, recording per-stage p50 and p99 alongside tokens in and tokens out. A lever’s effect is then a measured delta rather than an impression, and the same harness catches the regression when a model version moves under the feature. Profile by token distribution rather than request count: the document-summary path might send 4,000 input tokens a call where the quick-lookup path sends 300, and that distribution, not a feature’s share of the traffic, identifies what is dragging prefill time. Output tokens per second is the companion signal, computed as OutputTokenCount divided by InvocationLatency minus TimeToFirstToken. It separates a model generating more slowly from a model generating more tokens, which an end-to-end number alone cannot do.
With the diagnosis in hand, the order is clear. The p99 first-token wait is the fattest bar, so the input-side levers come first: trim retrieval from twenty passages to the three that rerank highest, cut conversation history to the turns that matter, and cache the stable prefix. Prompt caching is the clean win here, because the system prompt and instruction block repeat on every call in a session. Mark them with a cache checkpoint and they are not reprocessed each turn, so prefill time drops without touching the answer. Trimming context does double duty, shaving prefill time and often improving quality by dropping passages that were never relevant.
Streaming is the change that most improves how the feature feels, and it takes little work to add. It does not move the 4.2-second total at all, which is why a team measuring only averages will underrate it. It does move the first visible token from the end of the answer to the end of prefill, a few hundred milliseconds once the input-side cuts have landed, and for a chat feature that is the difference between responsive and broken. Ship it alongside the input-side cuts, not instead of measuring, because it masks a slow total rather than fixing one.
The model swap is the biggest lever and the one to reach for deliberately. Routing beats a blanket downgrade: send the short, factual questions to a fast small model and reserve the large one for the questions that need it, so the median gets much faster while the hard tail keeps its quality. Every such change has to be checked against an evaluation of answer quality, because a smaller model that is two seconds faster and wrong is not a win. Where the small model still is not fast enough and the quality floor rules out going smaller, latency-optimised inference runs the same model faster at a higher price per token, provided the model and Region are on the preview’s supported list.
The tail levers are for the p99 spikes that instrumentation ties to load rather than to any one request. If the slow tail correlates with peak traffic and requests are being throttled, a cross-Region inference profile spreads load across Regions under larger quotas. Provisioned Throughput instead reserves capacity for steady high-volume traffic, and the two are mutually exclusive. Neither helps a single uncontended slow request, so reach for them when the data shows contention rather than as a reflex. Parallelising the independent work is the smallest change on the list. Run the vector search and the metadata lookup at once instead of in series, and the faster of the two leaves the critical path for a code change and nothing else.
Worked example
The team instruments the pipeline and gets the picture above: p50 at 3.1 seconds and p99 at 11.8. The first-token wait dominates the tail at 4.5 seconds p99, generation adds 3.8, and a slow tool call adds another 2.6 when it fires. Retrieval, the stage two people wanted to rebuild, is 0.7 seconds at worst.
They attack the biggest bars in order. Reranking retrieval from twenty passages to four cuts roughly 3,000 tokens of input, and the first-token p99 falls from 4.5 to about 2.9 seconds. Marking the system prompt and instruction block as a cached prefix takes another slice off first-token wait on every turn after the first, dropping it to around 2.1. Capping max_tokens and prompting for a tighter answer pulls generation p99 from 3.8 to 2.7. The slow tool call turns out to be running serially after retrieval for no reason; moving it to run in parallel with the vector search removes it from the critical path except where the answer depends on its result mid-generation.
Then they add streaming, which changes none of those totals. It puts the first visible token at the end of prefill instead of the end of the answer, about 350 milliseconds at p50 once the input-side cuts have landed, so the feature feels responsive even on the slow path. The p99 end-to-end lands near 6 seconds, roughly half what it was, and the model is unchanged, so the benchmark set scores no worse than before. Only then, with the quality floor still respected, do they consider latency-optimised inference for the remaining prefill wait, and they leave the model swap on the shelf because they never needed it.
What’s worth remembering
- Measure p50 and p99 per stage. Averages hide the tail, and the tail is what users complain about.
- Prefill and decode scale differently. Prefill sets first-token wait and scales with input length; decode scales with model size and output tokens.
- Streaming changes feel, not total time. It moves the first visible token earlier, and Bedrock publishes
TimeToFirstTokenonly for streaming calls. - Cache the stable prefix. Prompt caching cuts prefill time, most when the stable prefix is large next to the variable part.
- Trim retrieved context and history. Less input shortens prefill and often improves quality too.
- Check every speed win against quality. A smaller model is the fastest lever but may answer worse; run an output evaluation.