Exam Room · Advanced Generative AI Developer

Preparing a Dataset for Fine-Tuning

· 31 min read

Generative AI Development · part of The Exam Room

The situation

A team is fine-tuning a foundation model on Amazon Bedrock to draft internal support replies in the house voice: a particular structure, a fixed sign-off, a calm tone, and a habit of quoting the ticket reference back to the customer. Prompt engineering got them most of the way. The system prompt has swollen to hundreds of words of tone instructions and worked examples, it costs tokens on every call, and the model still drifts out of voice about one reply in ten. Fine-tuning is the right tool for baking in behaviour that a prompt keeps having to re-teach.

They have three years of resolved tickets in a data warehouse: roughly forty thousand agent replies, of varying quality, written by dozens of people across several eras of the style guide. The instinct is to throw all forty thousand at the job and let scale sort it out. That instinct is the problem. A large fraction of those replies are off-voice, contradictory, or full of customer names, addresses, and card fragments. Some near-duplicate replies would land in both the training and the evaluation data. Feeding the model that heap teaches it the average of every era’s style, including the bad ones.

What they actually need is a deliberately curated set of examples that demonstrates the behaviour they want, in exactly the format the model expects, with nothing in it that leaks private data or inflates the evaluation score. The raw pile could not go in whole regardless. Bedrock’s default quota for the combined training and validation records in a fine-tuning job is ten thousand for most fine-tunable models, twenty thousand for the Nova family, adjustable through Service Quotas. Curation is the work, not a preliminary to it.

What actually matters

The first thing worth naming is what fine-tuning changes and what it does not. Fine-tuning on labelled examples teaches behaviour, format, and tone: how to respond, in what structure, with what voice. It does not reliably install fresh facts. If the model needs to know current pricing, this week’s policy, or a specific customer’s history, that knowledge belongs in retrieval at inference time, not baked into weights that go stale the moment they are trained. The clean mental split is that fine-tuning shapes how the model answers and retrieval supplies what it answers about; a serious support assistant usually needs both, fine-tuning for voice and structure, retrieval for the facts.

Given that, quality and consistency beat volume by a wide margin. AWS puts numbers on it in its own dataset guidance for Nova fine-tuning: two hundred samples minimum, two thousand to ten thousand recommended, quality prioritised over quantity. A few thousand clean, representative, identically formatted examples will produce a better tuned model than tens of thousands of noisy ones. The model fits the regularities in the set, so every inconsistency is a lesson too. If half the examples end with a sign-off and half do not, the sign-off turns up in roughly half the outputs. If the reference number is formatted three ways across the set, all three come back out at random. Inconsistent labels and formatting do not average out to something reasonable; they train the model to be inconsistent.

Consistency is not the same as sameness, though, and this is the tension to hold. The set has to be consistent in format and voice while still covering the real distribution of the task, including the edge cases. If every example is a simple happy-path cancellation, the model handles cancellations beautifully and falls apart on a billing dispute or an angry complaint. Coverage means the set spans the intents, tones, and awkward shapes the model will actually see in production, each rendered in the one consistent house format. AWS’s dataset guidance calls for the same spread: the full range of expected inputs, a mix of difficulty levels, and the edge cases. Representative of the real spread, uniform in presentation.

Then there is safety and hygiene, which is where careless datasets do real damage. Support text is dense with personal data: names, emails, addresses, phone numbers, card fragments, order histories. That has to be removed or redacted before anything is staged for training, because a model trained on it can regurgitate it, and because the raw data sitting in a bucket is a liability of its own. Duplicates and near-duplicates have to go, since they over-weight whatever pattern they repeat. The most damaging fault is leakage between the training split and the validation split: if a reply or a close paraphrase appears in both, the validation numbers look wonderful and mean nothing, because the model is being tested on what it memorised.

Finally, the mechanics. Bedrock fine-tuning takes the data as JSONL, one example per line, with fields matching the target model’s expected schema, staged in Amazon S3 for the job to read. The validation file is a second, separate S3 object you point the job at through validationDataConfig. It is optional, and Bedrock does not carve one out of the training file for you. The result then has to be judged on a held-out set the model never saw during training, because Loss curveThe plot of training error over time; the gap between the training and validation lines is how you spot memorising rather than learning. tells you the model fit the data, not that it does the job.

What we’ll filter on

  1. Task fit, is the goal behaviour, format, and tone (fine-tuning’s job) rather than fresh factual knowledge (retrieval’s job)?
  2. Example format, are the records labelled examples in JSONL, matching the target model’s expected schema, with a train and a validation split?
  3. Consistency, is every example formatted and labelled the same way, so the model learns one regular pattern rather than several conflicting ones?
  4. Coverage, does the set span the real task distribution and its edge cases rather than repeating the happy path?
  5. Safety and hygiene, is PII removed, are duplicates gone, and is the train/validation split free of leakage?
  6. Evaluation, is there a held-out set to measure the tuned model against, separate from anything it trained on?

The landscape

The pieces you assemble a fine-tuning set from, and what each one is for:

Labelled prompt-completion pairs. The core unit for fine-tuning that teaches a task: an input and the exact output you want the model to have produced. For the support case, the prompt is the ticket context and the completion is the ideal house-voice reply. The model fits the mapping from the one onto the other. This is supervised fine-tuning, one of three customisation methods Bedrock documents alongside reinforcement fine-tuning and distillation.

JSONL in the model’s schema. One JSON object per line, with the field names the target model expects. The shapes differ by family, so a set formatted for one model may need reshaping for another. Non-conversational families such as Llama 3.1 take a prompt field and a completion field. Conversational ones do not: Claude 3 Haiku takes a system string and a messages array of alternating user and assistant turns, and the Nova models take the Converse shape, keyed on schemaVersion set to bedrock-conversation-2024. Match the documented schema of the model you are tuning, not a generic template.

The train split. The bulk of the curated examples, the data the job actually learns from. This is where consistency and coverage have to be right, because everything in here is a lesson.

The validation split. A separate, smaller file the training job reads during tuning to track how the model generalises as it learns, so you can catch OverfittingWhen a model stops learning the general pattern in your data and starts memorising the individual examples. while the job runs. Bedrock writes validation loss per epoch into your output bucket alongside the training metrics. It is optional and it is yours to build: there is no automatic carve-out from the training file, and Nova 2.0 supervised fine-tuning does not use one during training at all. It must not overlap the training split.

The held-out evaluation set. Examples set aside before training and never shown to the job at all, kept to judge the finished model. This is distinct from the validation split, which the job sees during training; the held-out set is the last honest measurement you have. Keep it representative of production and never let a curated training example or its paraphrase drift into it.

Distillation prompts. A different customisation method with a different data shape. For distillation you supply prompts alone, and Bedrock generates responses from a larger teacher model and fine-tunes a smaller student on them. Reach for it to move an established behaviour onto a smaller model, not to teach a voice you can already demonstrate in examples. Continued pre-training on unlabelled domain text, which used to sit alongside fine-tuning as an option here, is no longer one of the customisation methods Bedrock documents, so do not plan a dataset around it.

Reinforcement fine-tuning data. Prompts paired with a reference answer, in JSONL, scored during training by a reward function you write as an AWS Lambda function rather than by matching a target completion. Bedrock takes the prompts in the OpenAI chat completion format and requires at least a hundred records. It suits goals a program can score; house voice is not one of them.

Evaluation

Side by side

Data shape Teaches Labelled? Volume vs quality Format Use it for
Prompt-completion pairs Behaviour, format, tone ✓ Quality wins JSONL, model schema Supervised fine-tuning
Train split The task itself ✓ Quality wins JSONL in S3 What the job learns from
Validation split Generalisation during training ✓ Small, clean JSONL in S3, optional Catching overfit mid-job
Held-out eval set Nothing (never trained on) ✓ Representative Kept aside Judging the tuned model
Distillation prompts Nothing directly ✗ Coverage helps Prompts in JSONL Moving a behaviour to a smaller model
RFT prompts and reference answers What a reward function scores ✓ Moderate JSONL, chat format Goals a program can score

Reading the table against the support goal: the job calls for labelled pairs in the model’s JSONL schema, split into train and validation with no overlap, plus a held-out set kept back for the final judgement. Distillation is the wrong tool here, because the team is not moving a behaviour it already has onto a smaller model. Reinforcement fine-tuning is wrong for a different reason: house voice is demonstrated in examples, and a Lambda reward function has nothing reliable to score it against.

The solution

Curating the pairs is where the real work sits, and it starts with throwing most of the raw data away. From forty thousand replies, the team should select a few thousand that genuinely exemplify the house voice, rewriting where a good reply has a formatting wart and dropping anything off-voice, contradictory, or thin. Every surviving example gets normalised to one format: the same structure, the same sign-off, the reference rendered one way. The prompt side has to be consistent too, carrying the same fields in the same order, because the model learns the shape of the input as much as the output. Curating down and normalising is worth more than any amount of extra volume.

Coverage is the counterweight that stops curation from collapsing into a monoculture. Before finalising, check the set against the real intent and tone distribution: cancellations, billing disputes, technical problems, complaints, simple thanks, the awkward multi-part message. Each should appear enough times to teach its pattern, each in the one house format. A set that is consistent but narrow tunes a model that answers fluently and wrongly on the first intent the training data left out.

Hygiene runs across the whole set before anything is staged. Redact or remove PII from both the prompt and completion sides. Use pattern-based detection for the obvious identifiers and a review pass for the rest. Amazon Comprehend will locate or redact PII entities across a collection of documents, in English or Spanish, if you want a managed pass over the corpus. Whichever route you take, no live personal data reaches the training bucket. De-duplicate, including near-duplicates that differ only in a name or a date, because repeated examples over-weight whatever pattern they carry. Then split into train and validation and check for leakage across the boundary, deduplicating across the split and not just within each file. The fastest way to a meaningless validation score is a reply and its paraphrase landing on opposite sides.

The format and staging are the mechanical finish. Write the examples as JSONL, one object per line, with field names matching the target model’s documented schema, since the shape differs by model family and a mismatched schema fails the job or trains on garbage. Put the training file and the validation file in S3, point the Bedrock fine-tuning job at them through trainingDataConfig and validationDataConfig, and grant the job’s service role read access to the bucket. The validation file is optional, and nothing splits the training file on your behalf, so if you want validation loss reported per epoch you build that split yourself and check it for leakage before the job starts.

Evaluation closes the loop and has to be honest. Judge the tuned model on the held-out set that never touched training. Compare its replies against the original prompt-engineered baseline on what the team set out to improve: voice adherence, structural correctness, the sign-off, the reference quote, and whether any reply states something the ticket context did not support. A drop in training loss is not success; success is the tuned model beating that baseline on held-out examples. If it does not, the answer is almost always in the data (coverage gaps, residual inconsistency, too few examples of a hard intent) rather than in the training settings.

Worked example

Start with a raw warehouse row, which is exactly what must not go into training:

Agent reply, ticket 44821:
"Hi Dana Whitfield, cancelled your Pro plan (card ending 4417,
dana.whitfield@example.com). - Marcus"

It is off-format (no house structure, no reference quoted back, an ad-hoc sign-off) and it is full of PII. Curated, normalised, and redacted, it becomes one clean prompt-completion pair. The completion carries the house voice; the prompt carries a consistent set of fields; the personal data is gone:

{"prompt": "Intent: cancellation\nPlan: Pro\nReference: 44821\nTone: neutral\nDraft a house-voice reply.", "completion": "Thanks for getting in touch. I've cancelled your Pro plan, effective at the end of your current billing period. Your reference for this is 44821. If there's anything else we can help with, just reply here.\n\nBest,\nThe Support Team"}

That record is the non-conversational shape, which is what Llama 3.1 takes. A conversational family carries the same content differently: Claude 3 Haiku takes a system string and a messages array, and Nova takes the Converse shape. Every other surviving example gets the same treatment: same prompt fields, same reply structure, same sign-off, reference always rendered the same way, no live names or card fragments anywhere. Then the set is split, with a leakage check so no reply and its paraphrase straddle the boundary. Both files go to S3 for the job to read:

From raw tickets to a fine-tuning job Raw tickets ~40,000 replies mixed voice, PII, duplicates Curate & clean select best examples, normalise format, redact PII, de-dupe most rows dropped on purpose Train split JSONL, model schema the bulk of examples Validation split no leakage from train watched during job Held-out eval never trained on final judgement Bedrock fine-tuning job reads from S3 Tuned model scored on eval set Train and validation feed the job; the held-out set stays out of training and scores the result. Fine-tuning teaches voice, structure, and tone; pair it with retrieval for facts that change.

The result is that the model no longer needs a hundreds-of-words tone preamble on every call, because the voice is in the weights. The prompt shrinks to the ticket context, the token count per call drops, and drift falls because the behaviour was trained rather than requested. Set against that, a custom model is not served from the shared on-demand pool: you either purchase Provisioned Throughput for it or create a custom model deployment for on-demand inference, which is a separate charge from the tokens. The facts the reply depends on still come from retrieval at inference time, because those change and fine-tuning would only freeze a stale copy.

What’s worth remembering

  1. Fine-tuning teaches behaviour, format, and tone, not fresh facts; when the answer depends on knowledge that changes, pair the tuned model with retrieval rather than baking the facts into weights.
  2. Quality and consistency beat volume; AWS’s own dataset guidance sets a floor of two hundred samples and recommends two thousand to ten thousand, with quality prioritised over quantity.
  3. Consistency is not sameness; the set must still cover the real task distribution and its edge cases, each rendered in the one house format.
  4. Build the validation split yourself, since Bedrock does not carve one out of the training file, and check for leakage across the boundary; a reply and its paraphrase on opposite sides makes the validation score meaningless.
  5. Judge the tuned model on a held-out set it never saw, on the metrics you set out to improve; falling training loss shows it fit the data, not that it does the job.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.