The situation
A team ships a Bedrock-backed assistant that answers customer questions from an internal knowledge base. It works, mostly. Then someone proposes swapping the underlying model for a cheaper one, and the room splits: half the team says quality will drop, half says it will hold, and nobody can settle it because there is no measurement. The last three prompt changes were merged on the strength of “looks better to me” and a couple of hand-picked examples that happened to be open in a tab.
The pattern repeats every time anything changes. A retrieval tweak that helps five questions someone remembers might be breaking fifty nobody has checked. A prompt edit that fixes a complaint about tone might have loosened a refusal that used to hold. Each change is argued from anecdote, and the anecdotes get picked after the change, so they support it. The team has no way to say, in a number that means the same thing this week as last, whether the assistant got better or worse.
What they are missing is a fixed set of inputs with agreed-upon right answers: a golden dataset. Everything downstream, the model choice, the prompt library, the RAG configuration, is only as measurable as this dataset is representative. A biased or thin eval set misses regressions and then reports them as passing scores, which is worse than having no number at all.
What actually matters
A golden dataset, also called a ground-truth set, is a collection of representative inputs each paired with an accepted answer or, where a single answer is too rigid, a set of acceptance criteria the response has to satisfy. It is the reference the system is measured against. The value of every evaluation you ever run is capped by how honestly this set reflects what real users actually send, so the properties that make it trustworthy matter more than any tooling choice that comes later.
The first property is representativeness of the real distribution. The dataset has to look like production traffic, not like the questions that are easy to write. That means the common, boring middle in roughly the proportion it actually occurs, plus deliberate coverage of the cases that break systems: edge cases, ambiguous phrasing, long inputs, adversarial inputs, and the known-hard questions the team already knows are shaky. Skew the set toward tidy questions and you get an eval that reports high scores while real users hit the failures the set never sampled.
The most-missed slice is the questions the system should not answer. A trustworthy assistant refuses when the knowledge base does not cover something, when the request is out of scope, or when answering would mean inventing a fact. If the golden dataset contains only answerable questions, you can never measure appropriate refusal, and a model that answers an unanswerable input anyway scores identically to one that declines. Known-unanswerable cases, with “should refuse” as their accepted answer, are what let you catch the failure mode that hurts users most.
Then there is provenance and bias in how the examples are sourced. Real traffic and logs give you the true distribution but need scrubbing and labelling. Subject-matter experts give you authoritative answers and can invent the rare, high-stakes cases logs have not seen yet. Synthetic generation, using a model to produce test inputs, is fast and fills gaps, but a set built only from synthetic data reproduces the generating model’s coverage gaps and phrasing habits. It then measures how you handle the questions a model produces rather than the ones a human types. Synthetic examples belong in the mix, not as the whole of it.
Labelling is where subjective quality gets pinned down. For clear-cut tasks the accepted answer is obvious, but for tone, helpfulness, and “is this answer actually correct and complete” you need human judgement, applied consistently by people who know the domain. The mechanism is a labelling workflow: route examples to human labellers or reviewers against a written rubric, so the labels are consistent rather than tracking one engineer’s mood. Amazon SageMaker Ground Truth managed exactly that workflow and still runs it for teams already on it, but AWS documents it as no longer open to new customers, with no new features planned. A fresh build staffs the workflow itself, with its own annotators or a vendor workforce.
Finally, the dataset is only reusable if it is disciplined over time. You need a held-out slice you never look at while tuning, or you will OverfittingWhen a model stops learning the general pattern in your data and starts memorising the individual examples. prompts to the eval and get an inflated number. And you need the whole thing versioned, so a score from this month and a score from last month are comparing the same yardstick rather than one that was edited in between.
What we’ll filter on
- Distribution match: does the set mirror real traffic, including the boring common cases in their real proportion?
- Hard and edge coverage: are known-hard, ambiguous, and adversarial inputs deliberately included?
- Refusal coverage: are known-unanswerable and out-of-scope questions present, labelled as “should refuse”?
- Sourcing balance: does it draw from real logs and experts, not synthetic generation alone?
- Label quality: is subjective quality labelled by domain humans against a consistent rubric?
- Reusability: is there a protected held-out split, and is the dataset versioned for comparable results over time?
The landscape
Sourcing from real traffic and logs. Mine production requests, or pre-launch pilot logs, for genuine inputs. This is the truest picture of the distribution and surfaces phrasings nobody would have invented. The cost is effort: logs need de-duplicating, scrubbing of personal data, and answers attached, because a raw request has no accepted answer until someone supplies one. It also cannot cover cases that have not happened yet, which is where experts and synthetic generation come in.
Subject-matter experts. Domain people write both the inputs and the authoritative answers, and they can author the rare, high-stakes cases that logs rarely contain: the compliance edge, the dangerous misunderstanding, the question that must be refused. Expert time is scarce, so use it on the hard and high-consequence slices rather than the common middle that logs already cover well.
Synthetic generation. Use a model to generate candidate inputs, paraphrases, and edge variations at volume. Good for widening coverage, probing with adversarial phrasings, and filling thin categories. The trap is building the set only this way: synthetic-only data carries the generator’s stylistic and topical bias, over-represents the phrasings it produces most often, and misses the awkward, misspelt, half-formed way real people actually write. Treat synthetic examples as a supplement that a human reviews, never as the ground truth itself.
Consistent human labelling. However inputs are sourced, the accepted answers and quality judgements for anything subjective need consistent human labelling: a written rubric, a workforce that applies it (your own team or a partner), and a review pass that catches disagreement. This is the mechanism that turns “we think this answer is good” into a repeatable label other people would agree with. The managed options here have thinned out. AWS documents both SageMaker Ground Truth and Amazon Augmented AI as closed to new customers, and Amazon Mechanical Turk, the public workforce behind them, closes permanently on 30 September 2026. A new build runs the workflow with its own annotators or a vendor workforce, and the rubric is the part that carries over.
Stratification and sizing. Rather than one undifferentiated pile, divide the set into strata: topic areas, difficulty bands, input lengths, answerable versus should-refuse. Sizing follows from wanting each stratum big enough that a change moving it is visible above noise, so a small but deliberately stratified set that covers every category beats a large set that is ninety per cent easy questions. Report scores per stratum, not just one blended average, because a headline number can hold steady while the refusal stratum collapses underneath it.
Held-out split and versioning. Partition the set into a development slice you iterate against and a held-out slice you check only occasionally, to catch prompts that have been overfit to the visible eval. And version the whole dataset, in source control or a data store, so every result is tagged with the dataset version that produced it and month-over-month comparisons are honest. An eval set edited in place, with no version, invalidates every historical number, and nothing in the scores shows that it has.
Bedrock evaluation jobs and an LLM-as-a-judgeUsing a second model, prompted with a rubric, to score another model’s output when there’s no exact answer to diff against. harness. The dataset is the input to whatever runs the scoring. Amazon Bedrock evaluations come in four shapes: programmatic model evaluation jobs that compute metrics, jobs that route responses to human workers, jobs that use a judge model to score one model’s responses with a second, and RAG evaluations that score a knowledge base retrieve-only or retrieve-and-generate. All four read a prompt dataset from S3 in JSON Lines format, one JSON object per line, with a default quota of 1,000 prompts per job. Alternatively you run your own harness where a strong model grades each response against the accepted answer or rubric. The harness is interchangeable; the dataset is the durable asset.
Evaluation
Side by side
| Source or practice | Distribution fidelity | Covers hard and refusal cases | Bias risk | Effort or cost | Best role |
|---|---|---|---|---|---|
| Real traffic and logs | ✓ | ✗ (only what happened) | Low | High (scrub, label) | The backbone of the set |
| Subject-matter experts | ✗ (curated) | ✓ | Low | High (expert time) | Rare, hard, high-stakes cases |
| Synthetic generation | ✗ | ✓ (breadth) | High if used alone | Low | Widen coverage, fill gaps |
| Rubric-driven human labelling | n/a | n/a | Low (consistent) | Medium | Consistent subjective labels |
| Stratification and sizing | ✓ | ✓ | Low | Low | Make every category visible |
| Held-out split | n/a | n/a | Low | Low | Catch overfitting to the eval |
| Versioning | n/a | n/a | Low | Low | Comparable results over time |
Reading the table: no single source builds the set. Logs give the distribution, experts and synthetic generation give the hard and refusal coverage logs lack, rubric-driven labelling makes the labels consistent, and stratification, a held-out split, and versioning are what make the result trustworthy and reusable rather than a one-off snapshot.
The solution
Start from real traffic and build the backbone from the true distribution. Pull a sample of production or pilot requests, scrub anything sensitive, and stratify what remains by topic and difficulty so you can see the shape of what users actually ask. This anchors the set in reality and stops the common failure of an eval made entirely of questions the team found interesting to write. Attach an accepted answer or acceptance criteria to each one; a logged input without a label is a test case with no pass condition. A Bedrock prompt dataset has named fields for this: prompt for the input, an optional referenceResponse for the accepted answer, and an optional category, which is how you tag the stratum and filter the report card by slice afterwards.
Layer in experts and synthetic generation to cover what logs cannot. Have domain experts author the rare, high-stakes cases and the known-unanswerable and out-of-scope questions with “should refuse” as the accepted answer, because that slice is how you measure whether the system declines instead of inventing. Use synthetic generation to multiply coverage: paraphrase real questions, generate adversarial variants, and fill thin strata. Keep a human in the loop reviewing synthetic examples, and never let the set tip to synthetic-only, or you measure the generator’s world rather than your users.
Label subjective quality consistently, and protect a held-out slice. For anything where “correct” is a judgement (tone, completeness, factual accuracy against the knowledge base), route examples through a labelling workflow with a written rubric, staffed by your own annotators or a partner workforce, so two labellers reach the same verdict. Then split off a held-out portion you do not look at while tuning prompts. Iterate against the development slice; check the held-out slice only now and then. When the two diverge, the visible eval has been overfit and its scores have stopped meaning anything general.
Version the dataset and feed it into a repeatable harness. Store the set under version control or in a data store with an explicit version tag, and record which version produced every score, so a comparison across months is honest rather than a comparison of two different yardsticks. Point it at Amazon Bedrock model evaluation or RAG evaluation jobs, or at your own LLM-as-a-judge harness that grades each response against the accepted answer. The scoring mechanism can change; the dataset is the durable asset, and a model swap becomes a measured decision instead of an argument.
Then put the set on a schedule. Continuous evaluation workflows run it nightly or on every merge, so scores arrive as a series rather than whenever somebody remembers to look. Regression testing for model outputs then scores every prompt edit, model swap, and retrieval change against the same fixed set before it ships, which is how you catch the fifty questions a tweak broke while fixing five. Those scores become automated quality gates for deployments only once a threshold and an owner exist: a number that separates pass from fail, and a named person who decides what happens when a release lands on the wrong side of it. Express that threshold as a delta against the incumbent rather than an absolute floor, because run-to-run variation on generative output makes a fixed floor either tight enough to block on noise or loose enough never to block on anything. The mechanics of the gate, where it runs and what happens when it blocks a good release, are their own piece of work: turning a golden-set score into a deployment gate.
Worked example
The team that could not agree on the cheaper model builds a golden set before touching anything. They pull 400 real questions from three months of pilot logs, scrub them, and stratify: 60 per cent common account and product questions in their real proportion, 20 per cent known-hard cases (multi-part questions, ambiguous phrasing, long inputs), and 20 per cent that should refuse (out-of-scope requests, questions the knowledge base does not cover, and a few adversarial “just make something up” inputs).
Experts label the accepted answers, and the should-refuse cases get “declines, and states no fact the knowledge base does not hold” as their pass condition. A shared written rubric keeps the quality labels consistent. They version the set as v1 and split off 80 questions as held-out.
Now the swap is a measurement. They run each model against the development slice through its own Bedrock evaluation job, because an automatic job scores one model at a time, and read the scores per stratum. The cheaper model matches on the common middle and the known-hard cases, within noise. On the should-refuse stratum it drops sharply, returning answers to questions it should decline and stating facts the knowledge base does not hold. The blended average barely moved, so a single number would have let the swap through; the per-stratum breakdown is what showed the drop. They confirm the pattern holds on the held-out slice, then keep the current model for the refusal-sensitive paths and move on with a decision nobody has to argue about again.
What’s worth remembering
- A golden dataset is representative inputs paired with accepted answers or acceptance criteria; every downstream evaluation is only as trustworthy as this set is honest.
- Match the real distribution, including the boring common cases in their real proportion, not just the questions that were easy to write.
- Include known-unanswerable and out-of-scope cases labelled “should refuse”, or you can never measure appropriate refusal, and a model that answers anyway scores the same as one that declines.
- Source from real logs and subject-matter experts as well as synthetic generation; a synthetic-only set reproduces the generating model’s coverage gaps and phrasing.
- Stratify and report per stratum; a blended average can hold steady while the refusal stratum collapses underneath it.
- Version the dataset and tag every result with the version that produced it, so comparisons over time are honest.