Exam Room · Advanced Generative AI Developer

Lab: Give a Bedrock Chatbot a Memory

· 10 min read

Generative AI Development · part of The Exam Room

This is one of the hands-on labs alongside these posts. You get a working base and build the missing piece. The full lab is in lab-04-chatbot-memory.zip.

Before your first lab, do the one-time, once-per-account setup: run the zip’s preflight.sh to confirm your account is ready, then deploy the lab reaper, a standing backstop that auto-deletes any lab you forget to tear down after 24 hours.

The scenario

A support chatbot answers each message well and retains nothing from the one before. Every request starts from a blank slate, because the model holds no state between invocations. To carry a conversation you resend the earlier turns each time, and that transcript has to live somewhere durable between requests. This lab stores it in DynamoDB, keyed by a session id.

What you’re given

A Lambda that calls Bedrock, a DynamoDB table keyed by session_id with a TTL that clears old conversations (DynamoDB deletes expired items within a few days of their timestamp), and an IAM policy granting bedrock:InvokeModel plus GetItem and PutItem on that one table. The current handler sends only the latest message, so nothing earlier reaches the model. The gap is the memory.

Lab 04 solution architecture A CloudFormation stack contains a DynamoDB history table keyed by session id with a TTL, a Lambda function, and an IAM execution role. The Lambda reads the transcript with GetItem, calls Nova Lite through Converse with the recent turns replayed, then writes the new turns back with PutItem. The model sits outside the stack in Amazon Bedrock, serverless and billed per token. The model receives the replayed turns and holds no state itself. CloudFormation stack: genai-lab-04 Amazon Bedrock serverless, billed per token History table one item per session_id TTL clears old sessions GetItem, the transcript PutItem, the new turns Lambda function handler.py load, converse, save Converse, the recent turns Nova Lite receives the replayed turns, holds no state itself Execution role InvokeModel, plus GetItem and PutItem on this one table

Your task

Wrap the model call in a load and a save. Read the history, add the new turn, call with the whole transcript, add the reply, write it back. Concretely: fetch this session’s item from DynamoDB before the call (the item keeps the message list as JSON in its messages attribute), append the new user turn, and send the model the recent turns of that transcript instead of the single prompt it sends today. When the answer comes back, append it as an assistant turn and put the updated list back into the table. Every entry stays in the Converse messages shape the handler already uses.

Two details in there are worth the extra lines. The read is strongly consistent (ConsistentRead=True on the get), because the turn you are replaying was written a second ago. DynamoDB reads are eventually consistent by default, so a plain get can return the item as it stood before that write, which reads back as a broken memory. ConsistentRead set to true returns the most recent committed version.

The cap takes more than a bare slice. Take the last MAX_TURNS entries, then keep dropping the leading turn until the window opens on a user message. Once the history is full, cutting the oldest turn off an odd-length list leaves an assistant message first, and the Amazon Nova request schema states that the first turn should always be the user turn. Trimming back to a user turn keeps the replay valid for as long as the session lives.

Deploy and prove it

cd lab-04-chatbot-memory
./scripts/deploy.sh
./scripts/test.sh
./scripts/teardown.sh

The test states a fact in one turn and asks for it back in the next, same session. Before you wire it, the second answer omits the fact. After, it returns it, and a fresh session id starts a separate transcript.

When you want the reference answer, deploy it without editing anything (SRC=solution ./scripts/deploy.sh), or unfold it here:

Show the answer
messages = _load_history(session_id)          # GetItem, ConsistentRead=True
messages.append({"role": "user", "content": [{"text": prompt}]})

response = _bedrock.converse(
    modelId=MODEL_ID,
    messages=_recent(messages),               # replay recent turns
    inferenceConfig={"maxTokens": 512, "temperature": 0.2},
)
answer = response["output"]["message"]["content"][0]["text"]
messages.append({"role": "assistant", "content": [{"text": answer}]})

_save_history(session_id, _recent(messages))  # back to DynamoDB
def _recent(messages):
    window = messages[-MAX_TURNS:]
    while window and window[0]["role"] != "user":
        window = window[1:]
    return window

The ideas worth keeping

  • A model is stateless. Conversation memory is something you build by replaying the transcript on every call. Any scenario where a chatbot has to carry earlier turns is a store-and-replay problem.
  • Short-term memory is the recent transcript, and it lives in a fast key-value store keyed by session (DynamoDB here; a cache like ElastiCache is the other common home). A TTL keeps it from accumulating.
  • Memory costs tokens. Every replayed turn is input you pay for on every call, and an unbounded transcript eventually overflows the context window. So you cap the turns or summarise the older ones. Replaying recent turns versus summarising the distant past is the line between short-term and long-term memory.
  • The transcript stays in the Converse message shape, opens on a user turn, and alternates roles. That is why you append the assistant reply after each turn, and why a turn cap trims back to a user turn rather than cutting wherever the slice lands.

What’s worth remembering

  1. A model call carries no memory; you create memory by replaying the conversation on each call.
  2. Store the transcript in a session-keyed store (DynamoDB or a cache) and give it a TTL so it expires.
  3. Keep every turn in the Converse messages shape, start the replay on a user turn, and alternate roles from there.
  4. Replayed history is input tokens on every call, so cap or summarise it; an unbounded transcript eventually exceeds the model’s context window and raises the token bill each time.
  5. Short-term memory is the recent transcript; long-term memory is what you keep by summarising or storing facts beyond the window.
  6. Scope memory by session id so one user’s conversation never bleeds into another’s.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.