The situation
A product team has shipped an AI reply drafter on Amazon Bedrock. It reads a customer support thread, pulls a few relevant help-centre articles through retrieval, and drafts a response the agent can edit and send. It has been live for two months, thousands of drafts a day, and the team has a thumbs up and thumbs down button under each draft that nobody quite trusts.
The numbers look fine and feel wrong. The thumbs-down rate is low, but agents keep rewriting drafts before sending, and a chunk of drafts get discarded entirely. Support leads have a folder of screenshots of bad answers, but nothing connects a screenshot back to the exact prompt, the retrieved articles, and the model version that produced it. When someone asks whether last month’s prompt tweak actually helped, the honest answer is that nobody can measure it.
The team is after something better than anecdote: capture what users are really telling them, find where the feature fails, and feed that back into changes that are validated before they ship. And they have just noticed that the drafts, the threads, and the corrections are full of customer names, order numbers, and the odd card fragment, so wherever this data goes, it has to be governed.
What actually matters
The instinct is to add more buttons. The thing that actually matters is what happens to a signal after it is collected, because a reaction that lands nowhere is worse than no reaction: it consumes attention and changes nothing.
Start with the signal itself. Explicit feedback, the thumbs up or down and any correction the agent types, is high-value and low-volume; people rate a fraction of interactions and they rate the extremes. Implicit feedback is the opposite, abundant and noisy: an agent editing the draft heavily, retrying with a reworded request, or abandoning the draft and writing from scratch all carry information about quality, and they are emitted on nearly every interaction without anyone opting in. A loop that leans only on the explicit signal is reading a biased sample; the implicit signals are what give you coverage.
A signal is only useful if you can trace it back to what produced it. A thumbs-down with no context is a number; a thumbs-down joined to the exact prompt, the retrieved ChunkingSplitting documents into retrievable pieces before embedding them – small enough to match precisely, big enough to still make sense., the model and version, and the final response is a debuggable case. That join is what the rest of the system is built on, and it splits cleanly into two stores: the model invocation log that Bedrock can capture for you, holding the request and response, and your own event store holding the user reaction keyed by the same interaction id. Neither half is useful without the other.
Once cases accumulate, the value is in the clusters, not the individual gripe. A single bad draft is noise; forty bad drafts that all involve refund policy, or all cite the same stale article, or all fail on threads in a particular language, are a diagnosis. Clustering the failures is what turns a screenshot folder into a prioritised list, and each cluster points at a different lever.
The levers matter because most feedback does not call for touching the model at all. A cluster caused by a stale or missing document is a retrieval fix; a cluster caused by wrong tone or format is a prompt or few-shot change; a cluster caused by a consistently unsafe answer is a guardrail rule. Customising the model is the heaviest lever and the last one to reach for, justified only when you have enough high-quality, well-labelled data that smaller changes cannot address. Reaching for a training job when a prompt edit would do is the classic over-correction.
And nothing ships on vibes. Every change, prompt, retrieval, or model, gets validated against a held-out evaluation set built from real failures before it reaches users, because a change that fixes one cluster routinely regresses another. The loop is only safe if the validation gate is real, and if you watch for feedback bias, since the users who rate are not the users who do not, and optimising to the raters degrades quality for everyone else without showing up in the ratings.
What we’ll filter on
- Signal coverage, do you capture both the explicit reactions and the implicit edit, retry, and abandonment signals?
- Traceability, can every signal be joined back to its exact prompt, retrieved context, response, and model version?
- Diagnosis, does the design surface failure clusters rather than isolated complaints?
- Lever fit, does the feedback route to the smallest change that fixes it, prompt or retrieval before fine-tuning?
- Validation, is every change measured against a held-out eval set before rollout?
- Governance, is feedback data treated as potentially containing PII and governed accordingly?
The landscape
Explicit feedback capture. The thumbs up or down, a star rating, or a free-text correction the agent submits alongside the draft. Quick to add, unambiguous in intent, and directly attributable to one interaction. The limits are volume and bias: only a slice of interactions get rated, and raters skew toward the strongly good and strongly bad. Corrections are the richest form here, because the edited text is close to a gold answer for that case.
Implicit feedback capture. Behavioural signals emitted without the user deciding to give feedback: how much the agent edits the draft before sending (edit distance), whether they retried with a reworded prompt, how long they dwelled, whether they abandoned the draft entirely. Abundant and unbiased by opt-in, but noisy and correlational, a heavy edit might mean a bad draft or a picky agent. Best used in aggregate and as a coverage layer over the sparse explicit signal.
Feedback interfaces and rating systems. Two halves of the same control, usually designed together and failing for different reasons. The feedback interface is what the end user is shown and how much effort a reaction takes: where the control sits, whether it interrupts sending the reply, and whether anything is captured besides the click. An interface that requires a written sentence before it records a rating gets better data from far fewer people. Rating systems for model outputs are the scale behind the control, and the scale sets the volume: a binary thumb collects heavily because it takes one click and no judgement, while a five-point scale collects almost nothing because a busy agent stops to work out whether a draft is a three or a four. A binary control with an optional follow-up box is the usual compromise, since a thumbs-down with no captured context is a count, not a signal.
Model invocation logging. Amazon Bedrock can log request and response data to Amazon S3, Amazon CloudWatch Logs, or both, including the prompt, the completion, and metadata such as the model id, the calling principal’s ARN, and input and output token counts. It covers calls made through the bedrock-runtime endpoint, so Converse, ConverseStream, InvokeModel, and InvokeModelWithResponseStream. It is off by default and configured per account per Region, and the destination bucket or log group has to sit in that same account and Region. Request and response bodies are logged inline up to 100KB; anything larger, and any binary data, is written as a separate object in S3 under the data prefix, so a trace that has to survive long threads needs an S3 destination rather than CloudWatch Logs alone.
Your own event store. A table or log you control, keyed by interaction id, holding the user reaction, the retrieved chunk ids, the app-side context, and the outcome. Amazon DynamoDB or an S3-based event log both work. This is where the explicit and implicit signals live, and where they join to the invocation log to make a complete, queryable case.
Failure clustering and analysis. Grouping the joined cases to find where the feature fails: by topic, by cited document, by language, by outcome. This can be as simple as querying the event store with Amazon Athena, or as involved as embedding the failed inputs and clustering them. The output is a ranked list of failure modes, each feeding either the eval set or a guardrail rule.
The evaluation set. A curated, held-out collection of real inputs with known-good outputs, drawn straight from the failure clusters. Amazon Bedrock evaluations score against this set three ways: programmatic metrics, a second model acting as judge, or a team of human workers. A custom prompt dataset is JSON Lines in S3, one object per line with a prompt and optional referenceResponse and category fields, and an automatic evaluation job takes up to 1,000 prompts. This is the gate: a change is only an improvement if it moves the score without regressing the rest.
Guardrails. Amazon Bedrock Guardrails enforce rules independent of the prompt: Denied topicsSubjects you describe in plain language that a Bedrock Guardrail refuses to discuss, whichever way a user phrases the request., content filters, word filters, Contextual grounding checkA Guardrail check that tests an answer against the documents it was given and flags claims the source doesn’t support., automated reasoning checks, and sensitive-information filters that block or mask PII. Those filters cover a built-in entity list (names, addresses, card numbers) plus custom regex patterns, which is how a house-specific format like an order number gets caught. Feedback that surfaces a consistent unsafe or off-limits answer becomes a guardrail rule, which is a faster and more reliable fix than hoping a prompt edit holds.
Annotation workflows and human labelling. The structured second pass, where a trained reviewer scores a sampled output against a written rubric instead of reacting to it in the moment, and where raw reactions become clean, labelled data that a training job can actually use. Reviewers score against the same rubric the automated judge runs, so the two stay comparable and the human scores can calibrate the judge, and the reviewed cases are where a reference set comes from. The two managed services that used to cover this are closed to new customers: Amazon SageMaker Ground Truth for labelling jobs and Amazon Augmented AI for routing low-confidence outputs to reviewers. Both continue for existing customers, and neither is getting new features, so a new build should not start on either. The supported managed route is now a Bedrock human evaluation job, which provides the reviewer portal and work team, capped at 50 workers per team, and this is one of the places a human genuinely belongs in the pipeline. Annotation consumes reviewer time, so the sample rate sets how deep the second pass can go.
Model customisation. The heavy lever. Bedrock offers three methods. Supervised fine-tuning trains on labelled input-output pairs, which is the shape an agent’s correction produces. Reinforcement fine-tuning takes prompts, generates several candidate responses each, and scores them with a reward function you write as a Lambda function or delegate to a judge model, training with Group Relative Policy Optimization; it accepts existing Bedrock invocation logs as its dataset, so the trace this loop already captures is a valid input. Distillation transfers behaviour from a larger teacher model to a smaller student. Support is per model and per Region rather than universal, so check the current list before planning around it. DPO-style preference training on preferred-versus-rejected pairs is not a Bedrock customisation method; it runs as a training job on SageMaker AI.
Evaluation
Side by side
| Option | Captures signal | Traces to context | Finds clusters | Feeds a change | Governs PII | When it fits |
|---|---|---|---|---|---|---|
| Explicit feedback | ✓ | ✗ | ✗ | ✓ | ✗ | High-value, low-volume ratings and corrections |
| Implicit feedback | ✓ | ✗ | ✗ | ✓ | ✗ | Broad coverage over the sparse explicit signal |
| Feedback interface / rating scale | ✓ | ✗ | ✗ | ✗ | ✗ | Setting how much a reaction costs the user |
| Model invocation logging | ✗ | ✓ | ✗ | ✗ | needs care | Recording what the model saw and said |
| Own event store | ✓ | ✓ | ✗ | ✗ | needs care | Joining reactions to invocations by id |
| Failure clustering | ✗ | ✗ | ✓ | ✓ | ✗ | Turning cases into a ranked list of failure modes |
| Evaluation set | ✗ | ✗ | ✗ | ✓ | ✗ | The gate every change passes before rollout |
| Guardrails | ✗ | ✗ | ✗ | ✓ | ✓ | Enforcing safety rules independent of the prompt |
| Annotation workflow (human review) | ✗ | ✗ | ✗ | ✓ | needs care | Clean human labels for training data |
| Model customisation | ✗ | ✗ | ✗ | ✓ | needs care | Enough labelled data to justify a training job |
No single row is the answer. The loop is capture (explicit and implicit) joined through logging and the event store, analysed into clusters, routed to the lightest fitting lever, and validated against the eval set, with governance sitting across every store that touches user text.
The solution
Capture both signals, and join them by id. Keep the thumbs up or down and the agent’s correction, they are the highest-value data you have, and the correction is nearly a gold answer for that case. But do not stop there, because ratings are sparse and skewed. Record the implicit signals too: edit distance between the draft and what was sent, whether the agent retried, whether the draft was abandoned. The single decision that makes any of it usable is a shared interaction id. Turn on Bedrock model invocation logging so the prompt, retrieved context, and response land in S3 or CloudWatch Logs, write the user reaction to your own event store keyed by that same id, and now every signal joins back to exactly what produced it. Carry the id in the request metadata on the invocation so it is captured in the log record rather than inferred from timestamps later. Without the join you have two piles of numbers; with it you have cases.
Cluster before you fix. Do not act on individual complaints. Query the joined data, in Athena over the S3 event log, or by embedding failed inputs and grouping them, to find where failures concentrate: a policy topic, a specific stale article, a language, an outcome. Each cluster is a diagnosis, and each points at a different lever. A cluster that all cites one outdated document is not a model problem. Sorting complaints into clusters is what stops the team from fine-tuning away a problem that a document update would have fixed.
Route to the lightest lever that fits. Most clusters resolve without touching the model. Wrong tone or missing format is a prompt or few-shot change. Stale or absent context is a retrieval fix, reindex the document, adjust chunking, tune what gets fetched. A consistently unsafe or off-limits answer becomes an Amazon Bedrock Guardrails rule, enforced independently of the prompt so it holds regardless of wording. Reach for model customisation only when you have accumulated enough high-quality, labelled data that smaller changes genuinely cannot close the gap. That is where a disciplined labelling workflow matters, your own annotators producing clean labels against a rubric, and where a Bedrock customisation job or a SageMaker AI training job runs. It is the last lever.
Validate every change against a held-out eval set. Build the evaluation set from the real failure clusters, inputs paired with known-good outputs, and hold it out. Before any change reaches users, score it with an Amazon Bedrock evaluation job, programmatic scoring for scale and human reviewers for the subtle cases, and compare against the current version. The gate exists to catch regressions. A prompt edit that fixes refund-policy drafts routinely breaks something else, and only a held-out set surfaces that. Watch feedback bias while you are at it. The agents who rate are not a random sample, so track quality on the whole population rather than on the interactions that got a thumbs down.
Govern the data end to end. Feedback text is some of the most sensitive data you hold, because it is verbatim user and customer content: names, order numbers, occasionally a card fragment. Treat every store, the invocation logs, the event store, and any training set, as containing PII. Guardrails sensitive-information filters mask PII in flight, and one detail here catches teams out: that masking does not reach the invocation logs. The logged input field holds the original, unmodified request whatever the guardrail did to the response, and content a guardrail blocked is written to the logs as plain text. Masking the logs is a separate control, a CloudWatch Logs data protection policy, which masks at ingestion for everyone without the logs:Unmask permission and does nothing for events already written. So apply that policy before the feedback loop starts collecting, lock down the logs and event store with least-privilege IAM, set retention so raw feedback does not accumulate forever, and never hand un-redacted user text to a labelling workflow or a training job.
Worked example
The team turns on the loop for a fortnight. Invocation logging is on, the thumbs and edit-distance signals write to a DynamoDB table keyed by interaction id, and an Athena query over the joined data ranks the failure modes.
The top cluster is unmistakable: drafts about refund eligibility get thumbs-down at four times the baseline rate, and even the ones sent are heavily edited. Pulling ten cases with their retrieved context shows the cause at once, every draft cites a help-centre article describing last year’s 14-day window; the policy changed to 30 days in the spring, but the old article is still the top retrieval hit. This is not a model failure. The draft accurately summarised a stale document.
The fix is a retrieval fix: update the article, reindex, and confirm the new version is what gets fetched. Before rolling out, the corrected inputs go into the eval set with known-good 30-day answers, and a Bedrock evaluation job scores the change. Refund-policy accuracy jumps and nothing else regresses, so it ships. A prompt rewrite would have papered over the symptom, and a training job would have taken weeks to fix a fact that belonged in the index. The loop pointed at the right lever because the signal was joined to the context that produced it.
A second, smaller cluster is different in kind: a handful of drafts offered goodwill credit the company does not provide, generated from threads where the customer was angry. Wrong tone is a prompt or few-shot job, but offering something that does not exist is a safety rule, so it becomes a Guardrails denied-topic entry that blocks the offer regardless of how the prompt is worded. Two clusters, two levers, one gate they both pass through.
What’s worth remembering
- Capture explicit and implicit feedback. Ratings and corrections are high-value but sparse and biased; edits, retries and abandonment give coverage.
- Join signals by interaction id. One id ties the reaction in your event store to the prompt, context and response in the invocation log.
- Cluster before you fix. Group failures by topic, document or language; clusters, not single complaints, give a ranked, diagnosable list.
- Use the lightest lever. Prompt or few-shot for tone and format, retrieval for stale context, a guardrail for safety, customisation last.
- Validate against a held-out set. Score every change with a Bedrock evaluation job before rollout, because a fix for one cluster routinely regresses another.
- Guardrails do not mask invocation logs. Use a CloudWatch Logs data protection policy, applied before collection starts.