The situation
A team runs a customer-support copilot on Amazon Bedrock, backed by a Claude model through the Converse API. A conversation is a list of messages; on every turn the whole list goes back to the model as input, because the model carries no state between calls. For a five-turn exchange this is invisible. For the long sessions the copilot actually gets, a subscriber working through a delivery problem across forty or fifty turns, it is neither invisible nor cheap.
Two things go wrong as the transcript grows. The bill climbs, because input tokens are charged on every call and the transcript is resent in full each time. Turn fifty is billed for turns one through forty-nine again, at the full input rate or, where prompt caching applies, at the lower cache-read rate. Eventually the accumulated history, the next user message and room for a reply exceed the model’s context window. Bedrock then returns a validation error, or the oldest turns have to be dropped mid-conversation, and the copilot can no longer cite the subscriber’s name or the ticket reference from twenty turns ago.
The team wants long conversations that stay coherent without the per-turn cost growing without bound. The problem underneath is which history actually has to travel on every call, and what to do with the history that does not.
What actually matters
The first thing to name is what a conversation costs. The model holds no state, so continuity comes from the application resending the transcript. Every message you keep in the list is input billed on this turn and every turn after it. Replay everything and the total is closer to proportional to the square of the conversation length, because each turn resends a transcript one turn longer than the last. The context window is the hard ceiling above that. It is a fixed number of tokens the model can attend to at once, and when the transcript approaches it there is no next turn to bill for.
So the design choice is how to trim history, and every option trades one of three things. The first is fidelity. The verbatim transcript is the highest-fidelity record there is; anything that shrinks it, a summary or a dropped turn, loses detail, and the detail it loses might be the one the subscriber needs three turns later. The second is what the compaction itself costs. Summarising is an extra model call, run periodically to rewrite old turns into something shorter, with its own input and output tokens and its own latency. It consumes tokens now to avoid more tokens later, which nets out on long sessions and is pure overhead on short ones. The third is durability. Some facts have to outlive the window and even the session, the subscriber’s address, that they are lactose-intolerant, the reference number of an open complaint, and those need a home outside the transcript entirely.
That third one draws the line. Short-term memory is the recent transcript, the last several turns kept verbatim so the model has the immediate thread. Long-term memory is everything durable, summaries of what came before and discrete facts pulled out and stored, retrieved and reinjected when relevant rather than carried on every call. AgentCore Memory is built around that split. It stores the turn-by-turn events of a session as short-term memory, and runs a background extraction over them that writes long-term records: semantic facts, user preferences, session summaries, and episodes. The strategies below are different ways of drawing the same boundary.
What we’ll filter on
- History sensitivity, does answering the next turn need the exact wording of old turns, or just their gist?
- Cost shape, does the per-turn input grow without bound, stay flat, or grow only slowly?
- Fidelity loss, how much detail does the strategy discard, and can a dropped detail break an answer?
- Compaction overhead, does the strategy add its own model calls or storage, and does the session run long enough for the saving to exceed it?
- Durability, do some facts need to survive past the window or past the session into the next one?
The landscape
Replay everything. Keep the full message list and resend it every turn. Highest possible fidelity, zero compaction machinery, and it is the right default for short conversations. Input tokens and latency then grow with every turn, the total grows faster still, and a long enough session runs into the context-window ceiling. This is the baseline the other strategies exist to fix, not a strategy to reach for on a copilot that gets long sessions.
Prompt caching on the replayed prefix. Resend the whole transcript, but have Bedrock serve the repeated prefix from cache rather than reprocessing it. Claude models with caching support on Bedrock offer both forms: implicit caching, which matches eligible prefixes without any request changes, and explicit caching, where you place a cachePoint in the system, tools or messages field of a Converse request. Tokens read from cache are billed at the model’s cache-read rate and do not count against the input-tokens-per-minute quota. The default cache lifetime is five minutes, with a one-hour option on current Claude models, and a changed prefix is a miss. This lowers the rate on replayed history without losing a word of it, and it moves the context-window ceiling not at all, because the cached prefix still occupies the window.
Sliding window of recent turns. Keep only the last N turns (or the last N tokens) and drop anything older before each call. Per-turn cost stops growing. It plateaus at whatever the window holds, which makes spend predictable and keeps you clear of the ceiling. What it loses is everything older: once a turn falls out of the window it is gone, so the ticket reference from turn three stops reaching the model the moment turn three ages out. A sliding window is simple and inexpensive, and correct only when old turns genuinely stop mattering.
Running summarisation. Periodically replace the older turns with a model-written summary. Keep the recent turns verbatim for immediacy, and every so often, when the transcript crosses a token threshold, make a separate model call that condenses the older block into a paragraph or two, then carry that summary in place of the raw turns. Cost grows slowly instead of linearly, because the old history travels as a short summary rather than a long transcript, and nothing falls off a cliff the way it does with a bare window. What it gives up is fidelity, and it adds a call: the summary is lossy by design, a detail can drop out or come back distorted, and each compaction consumes its own tokens and adds latency. This is the workhorse for long single sessions.
Fact extraction to a store. Pull durable facts out of the conversation and write them to a store, then retrieve the relevant ones on later turns instead of carrying the whole history. When the subscriber says they are lactose-intolerant or gives an address, extract that as a discrete fact, persist it (a database, or a vector store when you want to fetch facts by semantic relevance), and inject only the facts that matter to the current turn. This is long-term memory proper: facts survive the window, survive the session, and can inform a conversation weeks later. What it takes is machinery: something has to identify a fact, write it, and retrieve it, and the retrieval step can return the wrong facts or miss the right ones.
Managed conversational memory. Let the platform keep short-term and long-term memory for you. AgentCore Memory stores the session’s events, and what survives across sessions depends on the memory strategies attached to the memory resource. Specify no strategies and no long-term records are extracted at all, leaving short-term memory only. The built-in strategies run extraction and consolidation with predefined algorithms; built-in overrides keep that managed pipeline but let you replace its prompts; self-managed strategies hand you the extraction and consolidation yourself. Extraction is a background process, so a fact stated this turn becomes a long-term record some time afterwards rather than immediately. In every case you get the transcript-plus-summary split without writing the summarisation loop or running the store, around a reasoning loop that stays yours.
Most production copilots end up combining these rather than picking one: a sliding window for the immediate thread, running summarisation for the rest of the session, and fact extraction for the handful of things that must outlive it.
Evaluation
Side by side
| Strategy | Per-turn cost shape | Fidelity | Compaction overhead | Survives the window | Survives the session | Best for |
|---|---|---|---|---|---|---|
| Replay everything | Grows every turn | Full | None | ✗ | ✗ | Short conversations |
| Prompt caching on replay | Grows every turn, mostly at the cache-read rate | Full | None | ✗ | ✗ | Stable prefixes resent inside the cache lifetime |
| Sliding window | Flat (capped) | Recent only | None | ✗ | ✗ | Threads where old turns stop mattering |
| Running summarisation | Grows slowly | Lossy on old turns | Extra model call | ✓ | ✗ (unless persisted) | Long single sessions |
| Fact extraction to a store | Flat + retrieval | Exact on stored facts | Extraction + store + retrieval | ✓ | ✓ | Durable facts across sessions |
| AgentCore Memory | Managed | Summary-level | Handled by the platform | ✓ | ✓ | Cross-session recall without building it |
Cost shape and durability separate them. Replay, cached or not, is the one whose input grows every turn, and below a certain length that is fine because it is simplest and loses nothing. Caching lowers the rate on that growth and leaves the window ceiling where it is. Everything else caps or slows the growth by giving up some fidelity, and the two that reach across sessions do it by moving durable content out of the transcript entirely.
The solution
For the support copilot, layer the strategies, and let the layers follow how long each piece of information has to matter.
Keep a sliding window for the immediate thread. The last several turns carry the live back-and-forth, the thing the subscriber said two messages ago and the clarification you asked for, and those have to travel verbatim because their exact wording is what makes the next reply coherent. Size the window in tokens rather than turns, because a turn that pastes in an error log is worth ten short ones, and a token budget keeps the input predictable regardless of how chatty any single turn gets.
Add running summarisation for the rest of the session. When the transcript crosses a threshold, make a separate call that condenses everything older than the window into a short running summary, and carry that summary in place of the raw turns from then on. This is where fidelity is deliberately traded for room, so the summary prompt matters: instruct it to preserve concrete commitments, numbers, references, and unresolved questions, and to compress pleasantries and resolved detail. The overhead is an extra call with its own latency and tokens. Summarise on a threshold rather than every turn, and on a copilot that mostly gets short sessions, do not summarise at all.
Extract the durable facts to a store. A handful of things must outlive both the window and the session: the subscriber’s identity, delivery address, dietary constraints, the reference of an open complaint. Pull those out as discrete facts and persist them, and on later turns retrieve only the facts relevant to the current message rather than carrying all of them. A vector store fits here when you want to fetch facts by semantic relevance rather than by exact key, which is the same retrieval machinery behind a Bedrock RAG index, pointed at conversation-derived facts instead of documents. This is the layer that makes a returning subscriber feel known.
Turn prompt caching on underneath all of this. The stable part of each request, the system prompt, the tool definitions and the running summary, belongs ahead of the volatile recent turns, because a cache checkpoint only helps for content that does not change between calls.
If you would rather not build the summarisation loop and the store yourself, AgentCore Memory does both jobs around your own loop. Get the strategy set right: a memory resource with no strategies attached keeps the session events and extracts nothing durable, which looks like working memory right up until a subscriber comes back. Scope the namespaces by actor id so each subscriber’s extracted records stay isolated. You give up control over exactly what is kept, and you do not have to write and operate the compaction yourself.
Two mistakes are worth calling out. Reaching for summarisation on conversations that are never long enough to need it just adds a model call and latency for no saving; a plain sliding window, or even replaying everything, is cheaper and lossless below the length at which compaction saves more tokens than it consumes. And leaning on a bare sliding window for a copilot that needs continuity means the ticket reference leaves the input the moment that turn ages out, and the answer comes back without it. A window holds no long-term memory. If facts must persist, they have to be summarised or extracted, not merely windowed.
Worked example
A subscriber opens a session about a missed delivery. Turn three, they give the ticket reference and their address. Turns four through forty are back-and-forth: what happened, what the substitution policy is, whether they get a credit. By turn forty-five the raw transcript is large. Replaying it every turn is the biggest line on the bill and a slow crawl toward the context ceiling.
Under the layered approach, three things are happening at once. The sliding window holds roughly the last eight turns verbatim, so the model always has the immediate thread in full. Everything older has been folded, by a summarisation call that fired when the transcript first crossed the threshold and again later, into a short running summary: “Subscriber reports a missed delivery on the 28th, ticket GB-44821; agreed a credit for the missed box is being processed; substitution policy explained; subscriber still wants confirmation of the redelivery date.” And the two durable facts, ticket GB-44821 and the delivery address, were extracted to the store on the turn they were first mentioned, so they are retrievable exactly even if they never appear in the window or the summary again.
On turn forty-six, what the model receives is not fifty turns. It is the running summary, the ticket reference and address retrieved as facts, the last eight turns verbatim, and the new message. The input is a fraction of the full transcript, the cost per turn has stopped climbing, the window is nowhere near its ceiling, and the copilot still answers the redelivery question correctly because the reference it needs was preserved as a fact rather than left to age out of a window or blur inside a summary. When the subscriber comes back a week later about the same complaint, the stored summary and facts go into the first prompt of the new session, instead of it starting cold.
What’s worth remembering
- The model is stateless; continuity is the application resending the transcript, so every message you keep is input you pay for on this turn and every turn after it.
- Replaying the whole transcript makes per-turn cost grow with the conversation, and the total grows faster than length, until a long enough session hits the context-window ceiling.
- Short-term memory is the recent verbatim transcript; long-term memory is summaries and stored facts that are retrieved and reinjected rather than carried on every call.
- Size the recent window in tokens, not turns, so one log-pasting turn cannot blow the input budget and the per-turn cost stays predictable.
- Prompt caching lowers the rate on a resent prefix and leaves the context window exactly where it was, so it is a billing fix and not a compaction strategy.
- Match the strategy to session length and durability: replay or a plain window for short threads, summarisation for long single sessions, and extracted facts for anything that must outlive the session.