Exam Room · Advanced Generative AI Developer

Lab: Build RAG From Scratch

· 9 min read

Generative AI Development · part of The Exam Room

This is one of the hands-on labs that run alongside these posts. The scaffolding keeps fading: earlier labs were a line or two, this one hands you the whole pipeline except the retrieval step. The full lab is in lab-05-rag-from-scratch.zip.

Before your first lab, do the one-time, once-per-account setup: run the zip’s preflight.sh to confirm your account is ready, then deploy the lab reaper, a standing backstop that auto-deletes any lab you forget to tear down after 24 hours.

The scenario

You have used a Knowledge Base in the reading. Now write the retrieval step it runs for you. A support assistant has to answer questions about Greenbox, a fictional product that appears in no model’s training data, using only five short documents. With no managed vector store in the way, retrieval is the only moving part.

What you’re given

A Lambda that can call Bedrock for both embeddings and generation, the five documents in knowledge.py, the document-embedding step (done on cold start), the _embed() and _cosine() helpers, and the generation call that grounds the answer in whatever context it is handed. The one gap is retrieve().

Lab 05 solution architecture A CloudFormation stack contains a Lambda function and an IAM execution role scoped to bedrock:InvokeModel. The Lambda holds five documents, embeds them on cold start with Titan Text Embeddings V2, embeds each question the same way, ranks the documents by cosine similarity in memory, then has Nova Lite generate an answer from the top matches. Both models sit outside the stack in Amazon Bedrock, serverless and billed per token. CloudFormation stack: genai-lab-05 Amazon Bedrock serverless, billed per token A question, about Greenbox sources back Lambda function five documents baked in, embedded on cold start cosine ranking in memory embeds each document, then the question Titan Text Embeddings V2 generates the answer from the top matches Nova Lite Execution role bedrock:InvokeModel on foundation models

Your task

Turn a question into the best-matching documents. retrieve(query, k) has three moves: embed the query with _embed(), score that vector against every (doc, vector) pair in _DOC_VECTORS with _cosine(), and return the k highest-scoring document dicts, best first. A sort keyed on the score, descending, and a slice is all the ranking machinery it takes; four lines cover it.

That is the whole of retrieval: embed the query, score it against every document with Cosine similarityA measure of how closely two vectors point the same way, used as the default score for “how related is this text?”., take the top few. A vector store runs the same three steps, with an ANNIndex structures (HNSW graphs, IVF partitions) that answer the k-nearest-neighbours question fast by giving up guaranteed exactness – recall becomes a tunable knob rather than a certainty. index underneath so they still return in milliseconds at millions of documents instead of five.

Deploy and prove it

cd lab-05-rag-from-scratch
./scripts/deploy.sh
./scripts/test.sh
./scripts/teardown.sh

“When will my box arrive?” comes back with Thursdays and Fridays and names the delivery-days document as its source. “What is Greenbox’s carbon footprint per box?” comes back with “I do not know”, because no document supports an answer and the system instruction tells the model to answer only from the context it is handed.

When you want the reference answer, deploy it with SRC=solution ./scripts/deploy.sh, or unfold it here:

Show the answer
def retrieve(query, k=2):
    qv = _embed(query)
    scored = [(_cosine(qv, vec), doc) for doc, vec in _DOC_VECTORS]
    scored.sort(key=lambda pair: pair[0], reverse=True)
    return [doc for _, doc in scored[:k]]

The ideas worth keeping

  • Retrieval is embed, compare, rank. Every vector store, OpenSearch, pgvector, S3 Vectors, runs this same operation; the differences are speed, scale, and filtering, not the fundamental step. Building it by hand is why the vector-store choices click into place.
  • Embedding is a separate model from generation, with its own model id, its own price and its own Region list. Titan Text Embeddings V2 (amazon.titan-embed-text-v2:0) returns 1,024 dimensions by default, and 512 or 256 on request. Pick the wrong embedding model or the wrong distance metric and retrieval degrades before generation runs at all.
  • Grounding is two halves. Retrieve the right context, and instruct the model to answer only from it and to say when that context does not cover the question. Skip the second half and good retrieval still produces an answer drawn from training data; skip the first and there is nothing to ground on.
  • Naming the sources is one more field once you have retrieved, and it is what makes an answer auditable, the foundation of a citations-required assistant.

What’s worth remembering

  1. RAG retrieval is a cosine (or other distance) search over embeddings; a vector store makes it fast, it does not change what it is.
  2. Embedding and generation are different models with different ids, prices and Region lists; the embedding choice and its distance metric decide retrieval quality.
  3. Grounding needs both the retrieved context and a system instruction to use only that context and to say so when it does not cover the question.
  4. Returning the source ids alongside the answer is one more field in the response, and it makes the answer auditable.
  5. An “I do not know” when the context lacks the answer comes from the grounding instruction; drop that instruction and the model answers from training data instead.
  6. Once this is clear, a managed Knowledge Base is just this loop run at scale with sync, ChunkingSplitting documents into retrievable pieces before embedding them – small enough to match precisely, big enough to still make sense., and an index bolted on.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.