The whole GenAI track, sorted into the five scored domains. Each domain leads with the cheat sheet that anchors it, then the decisions to work through, the pop quizzes to drill, and the labs to build with your own hands. The domains overlap in real systems, so a few posts could sit under two headings; they are filed under the one they teach best.
How to use this
Read each domain heading and its scope line first, then run down the list. The order within each domain is deliberate: every list starts at the deciding-level idea and works toward its refinements, so the post above you is always the one the post below leans on. Top to bottom is the intended path. A line you can explain out loud, tick. A line that makes you hesitate is the next hour of revision. The pop quizzes are the fastest way to close a gap; the labs are the slowest and the ones that actually stick.
The five scored domains and their weight:
| Domain | Weight |
|---|---|
| 1. Foundation Model Integration, Data Management, and Compliance | 31% |
| 2. Implementation and Integration | 26% |
| 3. AI Safety, Security, and Governance | 20% |
| 4. Operational Efficiency and Optimization | 12% |
| 5. Testing, Validation, and Troubleshooting | 11% |
Before the first lab, do the one-time, once-per-account setup: run preflight.sh (it ships in every lab zip) to confirm your account is ready, then deploy the lab reaper once. The reaper is a standing backstop that auto-deletes any lab you forget to tear down after 24 hours, so a forgotten stack becomes a deleted stack instead of a running bill.
All in, the track is about 80 h 39 min of reading and roughly 11 h 30 min of hands-on lab work (lab times are active time; long-running jobs run on their own clock).
Weight is where the marks are, not where the difficulty is. Domain 1 is a third of the score on its own; give it a third of the time.
Domain 1: Foundation Model Integration, Data Management, and Compliance (31%)
About 32 h of reading and 5 h 30 min of hands-on lab time.
Picking the right foundation model, feeding it the right data, and grounding it in your own documents. The biggest domain by a wide margin, and the one where retrieval lives.
Anchor cheat sheets: Model Selection and Inference · RAG and Vector Stores · Model Customisation · 50 min
Background from the Service Manual: S3
Design the solution
About 1 h 13 min of reading.
- Review a workload against the AWS Well-Architected Framework and its Generative AI Lens before go-live, and turn the repeated answers into shared components.
-
Keep the model choice out of the application code, so switching model or provider is a configuration change rather than a deploy.
- ☐ Reviewing a GenAI Workload Against the Generative AI Lens · 36 min
- ☐ Switching Foundation Models Without Shipping Code · 37 min
Pick the model and the surface
About 5 h 35 min of reading.
- Decide whether GenAI is even the right tool for a problem, or whether a cheaper, deterministic approach wins.
- Choose a Bedrock model by matching capability, context window, latency, and cost to the workload, not by brand.
- Pick an inference surface (on-demand, provisioned throughput, batch) from the traffic shape, and know the hard limits that settle the SageMaker half: payload caps, processing timeouts, and Serverless being CPU-only.
- Separate model count from traffic shape: once it is dozens or hundreds of models, pack them onto shared infrastructure with inference components or a multi-model endpoint rather than an endpoint each.
- Say which serving surfaces a model can reach given where its weights came from, and what each one meters.
-
Say when a purpose-built AI service or a managed app like Amazon Quick beats building on a foundation model, and when an open-weight model or SageMaker JumpStart is the better host.
- ☐ Deciding Whether to Use GenAI at All · 30 min
- ☐ Choosing a Model From the Bedrock Catalogue · 44 min
- ☐ How to Pay for Serving a Model on Bedrock · 28 min
- ☐ Choosing an Inference Option for a GenAI Workload · 37 min
- ☐ Picking a Bedrock Model for High-Volume RAG · 28 min
- ☐ Choosing a Model for Code Generation · 26 min
- ☐ Open-Weight or Proprietary: Choosing How You Host a Model · 34 min
- ☐ SageMaker JumpStart or Bedrock for the Same Model · 28 min
- ☐ When a Purpose-Built AI Service Beats a Foundation Model · 28 min
- ☐ Buy or Build: Amazon Quick Versus a Custom RAG App · 26 min
- ☐ Choosing Between Kiro, Amazon Quick, and Bedrock · 26 min
Multimodal
About 2 h 05 min of reading.
- Pick the right Bedrock model for generating and for understanding images, audio, and video.
- Build cross-modal search with multimodal embeddings (a text query over images, and the reverse).
-
Wire a voice pipeline end to end: Transcribe in, Bedrock to reason, Polly out.
- ☐ Generating and Understanding Images, Audio, and Video on Bedrock · 39 min
- ☐ Searching Images and Text With Multimodal Embeddings · 27 min
- ☐ Building a Voice Assistant: Transcribe, Bedrock, and Polly · 34 min
- ☐ How to Build a Multi-Modal Bedrock Assistant for Insurance Claims · 25 min
Retrieval, embeddings, and knowledge bases
About 11 h 44 min of reading.
- Design a retrieval pipeline end to end: embedding model, distance metric matched to that model, vector store, index type, chunking, and top-K.
- Choose a vector store (OpenSearch, Aurora pgvector, S3 Vectors, a Bedrock Knowledge Base) from scale, latency, and cost, and name each one’s cost or latency floor.
- Know that Amazon Kendra is in maintenance mode and closed to new customers, and that a Bedrock managed knowledge base is the replacement.
- Enforce document-level ACLs with
userContexton retrieval, know it is optional and so fails open, and know Web Crawler is the connector it does not cover. - Use metadata filtering for freshness and multi-tenant access, and keep a knowledge base fresh without re-embedding everything.
- Reach for hybrid search and reranking when dense retrieval misses exact tokens, and parent-document retrieval when small chunks lose their context.
-
Route to text-to-SQL or structured extraction when the answer lives in a table rather than in prose.
- ☐ Building Permission-Safe Retrieval on a Bedrock Knowledge Base · 27 min
- ☐ Picking a Vector Store for Bedrock RAG · 23 min
- ☐ Which AWS Store Can Do Vector Search · 30 min
- ☐ Choosing a Vector Index: HNSW, IVF, and the Trade-Offs · 31 min
- ☐ Picking an Embedding Model for Retrieval · 22 min
- ☐ Choosing an Embedding Model for a Multilingual Corpus · 27 min
- ☐ Choosing an Embedding Dimension and Its Storage Cost · 25 min
- ☐ Choosing a Distance Metric for Embeddings · 25 min
- ☐ Dense, Sparse, or Hybrid Retrieval · 25 min
- ☐ Hybrid Search and Reranking for Bedrock RAG · 27 min
- ☐ Parent-Document Retrieval: Small Chunks, Big Context · 28 min
- ☐ How Many Chunks to Retrieve: Tuning Top-K · 27 min
- ☐ Metadata Filtering for Multi-Tenant Retrieval · 35 min
- ☐ Choosing a Chunking Strategy for Bedrock Knowledge Bases · 26 min
- ☐ Chunking Code, Tables, and Mixed Content · 27 min
- ☐ When a Document Won’t Fit the Context Window · 23 min
- ☐ Getting Documents Into a Bedrock Knowledge Base · 34 min
- ☐ Keeping a Knowledge Base Fresh Without Re-Embedding Everything · 32 min
- ☐ Building RAG When the Source Documents Change Daily · 24 min
- ☐ How to Build a Citations-Required RAG Over 50K Internal Documents · 28 min
- ☐ Grounding on Fresh Data: Tools or RAG · 27 min
- ☐ Agentic RAG: When Retrieval Needs to Reason · 35 min
- ☐ Extracting Structured Data From Documents at Scale · 33 min
- ☐ Retrieval Over Structured Data With Text-to-SQL · 28 min
- ☐ One Vector Index or Many · 35 min
Customise the model
About 3 h 12 min of reading.
- Choose between fine-tuning, continued pre-training, and distillation for a stated goal.
- Say when customisation beats retrieval and when it does not, and when the two belong together.
- Prepare and size a fine-tuning dataset, and read a training-versus-validation loss curve to spot overfitting or an under-trained run.
-
Import custom weights into Bedrock, and say how serving them bills compared with a Bedrock-native fine-tune.
- ☐ Fine-Tuning, Continued Pre-Training, or Distillation · 33 min
- ☐ Combining RAG and Fine-Tuning for a Legal Contract Assistant · 32 min
- ☐ Preparing a Dataset for Fine-Tuning · 28 min
- ☐ Tuning Fine-Tuning: Epochs, Learning Rate, and Batch Size · 33 min
- ☐ Importing Custom Weights into Bedrock · 27 min
- ☐ Promoting a Fine-Tuned Model into Production · 39 min
Data management and compliance
About 36 min of reading.
-
Pick the right AWS tool to profile, validate, and govern the data feeding a GenAI system, and say what each one catches.
-
☐ Picking the Right Tool to Check and Govern GenAI Data · 36 min
Prompt engineering and governance
About 4 h 35 min of reading.
The guide scores this material in Domain 1, under Task 1.6. It is listed again in Domain 2 beside the agent and delivery work it is built on, so a reader working either way round reaches it.
- Enforce role definitions and parameterised templates through Bedrock Prompt Management, with versions and an approval path.
- Test prompts for regression the way code is tested, and roll a bad version back.
- Reach for chain-of-thought and structured input patterns where the task needs reasoning shown.
-
Chain prompts in sequence with Bedrock Prompt Flows when the control flow is known ahead of time.
- ☐ Prompt Engineering Techniques That Move the Needle · 30 min
- ☐ Writing a System Prompt for a Production Assistant · 29 min
- ☐ Managing Prompts With Bedrock Prompt Management · 27 min
- ☐ How to Manage Prompts Across Thirty Services on Bedrock · 33 min
- ☐ Versioning and Rolling Back Prompts and Models · 29 min
- ☐ Tuning How a Model Samples: Temperature, Top-P, and Top-K · 25 min
- ☐ Building Deterministic Pipelines With Bedrock Flows · 34 min
- ☐ Summarising Long Conversations to Fit the Context Window · 29 min
- ☐ Handling Ambiguous Questions With Clarification · 39 min
Drill the pop quizzes
About 44 min of reading.
- ☐ Packing Many Models Onto One Endpoint · 4 min
- ☐ Hierarchical Chunking in One Line · 3 min
- ☐ Freshness and Access Are Metadata · 3 min
- ☐ What a Reranker Actually Fixes · 3 min
- ☐ Why Hybrid Search Finds ERR-4021 · 3 min
- ☐ The Silent Distance-Metric Bug · 3 min
- ☐ When Aurora pgvector Wins · 3 min
- ☐ OpenSearch Serverless’s Hidden Floor · 3 min
- ☐ S3 Vectors and the Latency Budget · 3 min
- ☐ Pop Quiz: Backoff or Breaker · 6 min
- ☐ Pop Quiz: What Model Registry Versions · 4 min
- ☐ Pop Quiz: Expand, Decompose, or Transform · 6 min
Build it
About 1 h 26 min of reading and 5 h 30 min of hands-on lab time.
Ready when you can pick a model and inference option for a stated workload, design a retrieval pipeline end to end (embedding model, distance metric, vector store, chunking, top-K, metadata), and say when fine-tuning beats retrieval and when it does not.
- ☐ Lab 05: Build RAG From Scratch · 8 min read + ~1 h hands-on
- ☐ Lab 11: Stand Up a Bedrock Knowledge Base · 18 min read + ~1 h hands-on
- ☐ Lab 07: Build a Data-Quality Gate · 9 min read + ~45 min hands-on
- ☐ Lab 08: Answer a Metric Question With Text-to-SQL · 10 min read + ~1 h hands-on
- ☐ Lab 12: Fine-Tune a Model and Read the Loss Curves · 20 min read + ~1 h hands-on
- ☐ Lab 13: Generate the Weekly Box Art · 21 min read + ~45 min hands-on
Domain 2: Implementation and Integration (26%)
About 19 h 01 min of reading and 3 h of hands-on lab time.
Turning a model into an application: prompts, tools, agents, memory, and how responses reach the caller.
Anchor cheat sheets: Prompt Engineering · Agents and Orchestration · 21 min
Prompts and sampling
About 2 h 53 min of reading.
The guide scores prompt engineering in Domain 1 (Task 1.6); it is repeated here because in practice it sits beside the agent and delivery work below.
- Apply the prompt techniques that actually change output quality (structure, few-shot, role, and where step-by-step reasoning helps).
- Write a production system prompt, and treat prompts as versioned, rollback-able assets through Bedrock Prompt Management.
-
Set temperature, top-P, and top-K deliberately for the task, from deterministic extraction to open-ended generation.
- ☐ Prompt Engineering Techniques That Move the Needle · 30 min
- ☐ Writing a System Prompt for a Production Assistant · 29 min
- ☐ Managing Prompts With Bedrock Prompt Management · 27 min
- ☐ How to Manage Prompts Across Thirty Services on Bedrock · 33 min
- ☐ Versioning and Rolling Back Prompts and Models · 29 min
- ☐ Tuning How a Model Samples: Temperature, Top-P, and Top-K · 25 min
Tools, agents, and orchestration
About 5 h 31 min of reading.
- Wire function calling through the Converse API and publish tools through an AgentCore gateway, and design tool schemas that are safe to expose to a model.
- Choose between an agent, Step Functions, and a Bedrock Flow for a given job, trading autonomy against determinism.
- Say what AgentCore adds for running agents in production, where the managed line sits between its harness and a loop you own, and how to coordinate more than one agent.
-
Choose a framework for a code-defined agent on the runtime, on whether it emits OpenTelemetry spans and consumes MCP natively rather than on loop syntax.
- ☐ How to Wire Function Calling Through Bedrock · 30 min
- ☐ How to Wire an LLM to Side-Effecting Actions with Bedrock AgentCore · 24 min
- ☐ Designing Safe Tool Schemas for an AgentCore Gateway · 40 min
- ☐ Orchestrating Multiple Bedrock Agents · 36 min
- ☐ Running Agents in Production With Bedrock AgentCore · 31 min
- ☐ Choosing an Agent Framework for the AgentCore Runtime · 36 min
- ☐ When to Orchestrate With Step Functions Instead of an Agent · 28 min
- ☐ Building Deterministic Pipelines With Bedrock Flows · 34 min
- ☐ Choosing Where an MCP Server Runs · 36 min
- ☐ Putting Brakes on an Autonomous Agent · 36 min
Memory and conversation state
About 1 h 30 min of reading.
- Design short-term and long-term memory for a chat assistant, and say what each is for.
- Choose where conversation state lives from the access pattern and retention need.
-
Keep a long conversation inside the context window by summarising older turns.
- ☐ Designing Short-Term and Long-Term Memory for a Bedrock Chat Assistant · 28 min
- ☐ Choosing Where to Store Conversation State · 33 min
- ☐ Summarising Long Conversations to Fit the Context Window · 29 min
Delivery and interaction
About 4 h 37 min of reading.
- Choose sync, async, or streaming delivery from the latency budget and the payload shape.
- Build an event-driven pipeline that processes documents asynchronously.
-
Handle an ambiguous question by asking for clarification instead of guessing.
- ☐ Delivering Responses: Sync, Async, or Streaming · 30 min
- ☐ Event-Driven GenAI: Processing Documents Asynchronously · 33 min
- ☐ Handling Ambiguous Questions With Clarification · 39 min
- ☐ Building a Deployment Pipeline for a GenAI Feature · 36 min
- ☐ Putting a GenAI Gateway in Front of Bedrock · 39 min
- ☐ Wiring a GenAI Assistant Into Systems You Cannot Change · 33 min
- ☐ Serving a GenAI Feature When the Data Cannot Leave · 33 min
- ☐ Cheat Sheet: Integration and Deployment · 34 min
Filed under another domain
About 2 h 52 min of reading.
Five posts are listed elsewhere, under the domain they teach best, and each one still carries work this domain leans on. They are not extra reading; they are the ones to re-read if this domain feels thin when you revise it on its own.
- Choosing Between Kiro, Amazon Quick, and Bedrock is the track’s Amazon Q Developer post, and where a finished assistant beats anything you assemble. Listed in Domain 1 beside the other surface choices. · 26 min
- Where Humans Belong in a GenAI Pipeline builds the review and approval step as orchestration, with Step Functions and Amazon Augmented AI (A2I). Listed in Domain 3 for the governance angle. · 33 min
- Routing Requests Between a Cheap and a Capable Model puts a classifier in front of two models and decides where a misroute is caught. Listed in Domain 4 as a cost lever. · 35 min
- Handling Throttling and Rate Limits Gracefully is the backoff, queue, and quota work that lands in the calling application. Listed in Domain 4 with the resilience material. · 34 min
- Tracing an Agent’s Decisions in Production wires X-Ray spans and CloudWatch Logs Insights across an agent’s tool calls, so a run can be reconstructed. Listed in Domain 4 with observability. · 44 min
Drill the pop quizzes
About 34 min of reading.
- ☐ Letting an LLM Take Actions · 4 min
- ☐ Turning Down the Randomness · 3 min
- ☐ Prompts Are Versioned Assets · 3 min
- ☐ Pop Quiz: Which Amazon Assistant · 6 min
- ☐ Pop Quiz: Strands, Agent Squad, or AgentCore · 5 min
- ☐ Pop Quiz: Where the MCP Server Lives · 6 min
- ☐ Pop Quiz: Getting a GenAI Feature in Front of Users · 7 min
Build it
About 43 min of reading and 3 h of hands-on lab time.
Ready when you can wire a tool an agent can call safely, choose between an agent, Step Functions, and a Flow for a given job, and pick where conversation state lives.
- ☐ Lab 01: Invoke a Foundation Model From Lambda · 13 min read + ~30 min hands-on
- ☐ Lab 06: Wire a Tool the Model Can Call · 9 min read + ~1 h hands-on
- ☐ Lab 03: Get Structured JSON Out With Tool Use · 11 min read + ~45 min hands-on
- ☐ Lab 04: Give a Bedrock Chatbot a Memory · 10 min read + ~45 min hands-on
Domain 3: AI Safety, Security, and Governance (20%)
About 12 h 43 min of reading and 30 min of hands-on lab time.
Keeping the app safe to run: guardrails, injection and exfiltration defence, identity and encryption, responsible-AI evidence, and the audit trail.
Anchor cheat sheet: Security and Responsible AI · 20 min
Scope the workload first
About 35 min of reading.
- Place a generative-AI use in the Generative AI Security Scoping Matrix, from a consumer app at Scope 1 through to a model trained from scratch at Scope 5.
- Say which controls the scope makes yours to implement and which you can only ask a provider to assure, and know that an application built on Bedrock is Scope 3 however managed the service is.
-
Name what fine-tuning adds when it moves a use to Scope 4: training-data provenance, the dataset as a governed artefact, and a custom model that can memorise what it was trained on.
- ☐ Deciding Which GenAI Security Controls Are Yours to Own · 35 min
Guardrails and moderation
About 1 h 47 min of reading.
- Configure a Bedrock Guardrail for PII, denied topics, and contextual grounding.
- Choose a guardrail strategy (managed, custom, or both) for a stated requirement.
-
Pick the right moderation service for the content type: Rekognition for images, Comprehend for text, Guardrails at the model boundary.
- ☐ Configuring Bedrock Guardrails for PII, Topics, and Grounding · 35 min
- ☐ Choosing a Guardrail Strategy: Managed, Custom, or Both · 35 min
- ☐ Content Moderation With Rekognition, Comprehend, and Guardrails · 33 min
- ☐ Pop Quiz: The Rule That Has to Be Provably Followed · 4 min
Attacks and data protection
About 3 h 58 min of reading.
- Defend against both direct prompt injection and the indirect kind that arrives through retrieved documents.
- Prevent data exfiltration through the model, and keep PII out of prompts and logs.
-
Red-team a Bedrock app to find these weaknesses before an attacker does.
- ☐ Defending a Bedrock App Against Prompt Injection · 29 min
- ☐ Defending Against Indirect Prompt Injection in RAG · 32 min
- ☐ Preventing Data Exfiltration Through an LLM · 34 min
- ☐ Red-Teaming a Bedrock Application · 29 min
- ☐ Keeping PII Out of LLM Prompts and Logs · 32 min
- ☐ Detecting Misuse of a Public GenAI Assistant · 36 min
- ☐ Deleting a Subscriber’s Data From a RAG System · 40 min
- ☐ Pop Quiz: Naming the LLM Risk in a Pen-Test Finding · 6 min
Identity, network, and encryption
About 1 h 50 min of reading.
- Secure a Bedrock app with least-privilege IAM and PrivateLink, so nothing reaches the model over the public internet.
- Encrypt every data surface end to end with KMS, using customer-managed keys where the requirement calls for them.
-
Separate inbound authorisation from outbound credentials for an agent, and match the flow (2LO, 3LO, on-behalf-of exchange, API key) to whose authority each downstream call needs.
- ☐ Securing a Bedrock App: IAM, PrivateLink, and Keys · 42 min
- ☐ Encrypting a Bedrock App End to End With KMS · 30 min
- ☐ Giving an Agent Credentials Without a Standing Key · 34 min
- ☐ Pop Quiz: Proving the Knowledge-Base Bucket Is Not Shared · 4 min
Responsible AI and governance
About 3 h 35 min of reading.
- Say which service produces which evidence: Bedrock evaluation jobs or fmeval for bias (Clarify, in maintenance since June 2026, for teams already on it), model cards for explainability, Audit Manager and invocation logging for the audit trail.
- Make a Bedrock app audit-ready, and govern model access across many teams.
-
Prove where AI content came from, and design where a human belongs in the pipeline and how the bot escalates to one.
- ☐ Checking a Bedrock Feature for Bias and Explainability · 31 min
- ☐ Making a Bedrock App Audit-Ready · 31 min
- ☐ Governing Model Access Across Many Teams · 39 min
- ☐ Proving Where AI Content Came From · 39 min
- ☐ Where Humans Belong in a GenAI Pipeline · 33 min
- ☐ Designing a Bot-to-Human Escalation Path · 35 min
- ☐ Pop Quiz: The Bias Number That Went Stale · 7 min
Drill the pop quizzes
About 27 min of reading.
- ☐ Keeping PII Out of Prompts and Logs · 4 min
- ☐ What Guardrails Enforce · 3 min
- ☐ The Eight Responsible-AI Dimensions · 5 min
- ☐ Measuring Bias With fmeval · 3 min
- ☐ LLM Explainability Is Traceability · 3 min
- ☐ The Prompt-and-Completion Record · 3 min
- ☐ Who Changed It vs What It Said · 3 min
- ☐ Turning Logs Into an Audit · 3 min
Build it
About 11 min of reading and 30 min of hands-on lab time.
Ready when you can name where access control belongs (retrieval and tools, not the prompt), attach a versioned guardrail, encrypt every data surface with KMS, and say which service produces bias or audit evidence.
- ☐ Lab 02: Put a Guardrail in Front of a Bedrock Model · 11 min read + ~30 min hands-on
Domain 4: Operational Efficiency and Optimization for GenAI Applications (12%)
About 13 h 46 min of reading.
Running it affordably and reliably: cost levers, caching, throughput, latency, resilience, and observability.
Anchor cheat sheet: Evaluation, Cost, and Operations · 17 min
Cost and caching
About 5 h 08 min of reading.
- Reach for the right cost lever for a given spend shape: caching, batching, request routing, model choice, or provisioned throughput.
- Tell prompt caching from response caching, and cache responses without serving stale answers.
- Attribute and cap spend with tagging, budgets, and quotas.
-
Route between a cheap and a capable model instead of paying top-tier rates for every request.
- ☐ How to Cut a Bedrock Bill Without Hurting Quality · 28 min
- ☐ Cutting Cost per Query in a RAG System · 35 min
- ☐ Cutting Ingestion Cost by Caching and Batching Embeddings · 35 min
- ☐ Cost Attribution and Tagging for GenAI Workloads · 40 min
- ☐ Cost Guardrails: Budgets, Quotas, and Model Choice · 39 min
- ☐ Budgeting Tokens for a Long-Document Workload · 31 min
- ☐ Routing Requests Between a Cheap and a Capable Model · 35 min
- ☐ Prompt Caching Versus Response Caching on Bedrock · 29 min
- ☐ Caching LLM Responses Without Stale Answers · 26 min
- ☐ Pop Quiz: A Bedrock Bill That Doubled Overnight · 7 min
- ☐ Flash Card: AWS Cost Anomaly Detection · 3 min
Throughput, latency, and resilience
About 4 h 39 min of reading.
- Choose provisioned throughput versus on-demand, and right-size the model units for a custom model.
- Cut end-to-end latency, and first-token latency with streaming, without hurting quality.
-
Build resilience: cross-region inference profiles, multi-region failover, graceful throttling, and a plan for a model deprecation.
- ☐ How to Match Bedrock Pricing to Workload Rhythm · 25 min
- ☐ Right-Sizing Provisioned Throughput for a Custom Model · 30 min
- ☐ Reducing End-to-End Latency in a GenAI App · 35 min
- ☐ Streaming Responses to Cut First-Token Latency · 26 min
- ☐ Multi-Region Resilience for a GenAI Service · 28 min
- ☐ Spreading Bedrock Load with Cross-Region Inference Profiles · 26 min
- ☐ Handling Throttling and Rate Limits Gracefully · 34 min
- ☐ Surviving a Model Deprecation on Bedrock · 29 min
- ☐ Auto-Scaling a Model Endpoint for Bursty GenAI Traffic · 39 min
- ☐ Pop Quiz: What to Scale a Model Endpoint On · 7 min
Observability
About 3 h 32 min of reading.
- Say what to watch in CloudWatch for a production Bedrock app, and which metrics warn you first.
- Trace an agent’s decisions to see why it did what it did, not just what it returned.
-
Separate the two logging features: model invocation logging watches inference, knowledge base logging watches ingestion, and enabling one gives you nothing of the other.
- ☐ Monitoring a Production Bedrock App · 43 min
- ☐ Tracing an Agent’s Decisions in Production · 44 min
- ☐ Finding the Documents That Never Reached the Knowledge Base · 29 min
- ☐ Keeping a Vector Store Healthy in Production · 40 min
- ☐ Dashboards for a GenAI Feature: Operations, Quality, and Business · 32 min
- ☐ Pop Quiz: The Agent Called the Wrong Tool, or the Tool Failed · 6 min
- ☐ Pop Quiz: Retrieval Got Slow and the Answers Got Worse · 9 min
- ☐ Flash Card: Amazon Managed Grafana · 4 min
- ☐ Flash Card: Amazon CloudWatch Synthetics · 5 min
Drill the pop quizzes
About 10 min of reading.
Ready when you can reach for the right cost lever for a given spend shape (caching, batching, routing, provisioned throughput), keep latency down without hurting quality, and say what to watch in CloudWatch and an agent trace.
- ☐ Cutting the Bill Without Losing Quality · 3 min
- ☐ Provisioned Throughput vs On-Demand · 4 min
- ☐ Seeing Into a Production Bedrock App · 3 min
Domain 5: Testing, Validation, and Troubleshooting (11%)
About 10 h 53 min of reading and 2 h 30 min of hands-on lab time.
Proving it works and finding out why it does not: evaluation, judging, hallucination measurement, A/B testing, and reproducibility.
Anchor cheat sheets: Evaluation Metrics · Evaluation, Cost, and Operations (the evaluation half) · 36 min
Evaluate and troubleshoot
About 8 h 54 min of reading.
- Pick the metric from the positive class and the error that costs: precision when a false alarm hurts, recall when a miss hurts, F1 when they cost about the same, and never accuracy on a rare positive class.
- Build a golden dataset, and score retrieval and generation separately rather than as one number.
- Design an LLM-as-a-judge rubric you can trust, and tell faithfulness from correctness.
- Measure hallucination, and trace a wrong answer back to the stage that caused it (retrieval or generation).
-
A/B test prompts and models in production, and make a given output reproducible.
- ☐ Picking an Evaluation Metric From the Cost of Being Wrong · 37 min
- ☐ Evaluating LLM Output With Bedrock Eval Jobs · 25 min
- ☐ Evaluating a RAG Pipeline End to End · 30 min
- ☐ Building a Golden Dataset for LLM Evaluation · 30 min
- ☐ LLM-as-a-Judge: Designing a Rubric You Can Trust · 27 min
- ☐ Measuring Hallucination in a RAG System · 30 min
- ☐ Why Your RAG Returns the Wrong Chunk · 41 min
- ☐ A/B Testing Prompts and Models in Production · 34 min
- ☐ Building a Feedback Loop From Users to Model Improvement · 34 min
- ☐ Making an LLM Output Reproducible · 32 min
- ☐ Evaluating an Agent’s Run, Not Just Its Answer · 36 min
- ☐ Turning a Golden-Set Score Into a Deployment Gate · 36 min
- ☐ Catching a Regression After the Deploy, Not From the Complaints · 38 min
- ☐ Which Bedrock Errors to Retry and Which to Surface · 34 min
- ☐ Finding Out Why a Prompt Stopped Behaving · 35 min
- ☐ Getting Evaluation Results in Front of the People Who Decide · 35 min
Drill the pop quizzes
About 29 min of reading.
- ☐ Evaluating Both Halves of RAG · 3 min
- ☐ Faithful but Wrong · 3 min
- ☐ LLM-as-a-Judge, and the Catch · 3 min
- ☐ Pop Quiz: Right Answer, Wrong Route · 4 min
- ☐ Pop Quiz: Reading a Bedrock Exception · 5 min
- ☐ Pop Quiz: The Gate That Blocks on Noise · 5 min
- ☐ Flash Card: Amazon Bedrock Model Evaluations · 6 min
Build it
About 54 min of reading and 2 h 30 min of hands-on lab time.
Ready when you can build a golden set, score retrieval and generation separately, tell faithfulness from correctness, and trace a wrong answer back to the stage that caused it.
- ☐ Lab 09: Evaluate the Pipeline · 9 min read + ~1 h hands-on
Tie it together
Two things pull all five domains into one picture. Read the first, do the second.
- ☐ Taking a GenAI Feature From Proof of Concept to Production maps the gap between a demo and a shipped feature across evaluation, safety, security, reliability, cost, observability, governance, and operations. It is the whole checklist in prose. · 38 min
- ☐ Lab 10: The Capstone builds one feature end to end and exercises every domain at once. If you can finish it without notes, you are ready. · 7 min read + ~1 h 30 min hands-on
Good luck. 🍀