Exam Room · AI Business Strategist

The Handbook the Assistant Never Read

· 24 min read

AI for the Business · part of The Exam Room

The situation

A general insurer with about 2,600 staff writes motor, home and small commercial cover. Fourteen months ago it put an internal assistant on Amazon Bedrock in front of underwriters, claims handlers and the contact centre. Around 900 people use it, and it answers roughly 11,000 questions a week: the referral threshold for a flood-zone property, whether a handler can settle at a given value without sign-off, what a wording says about escape of water.

The complaints come in three shapes. It quotes a flood referral threshold that was replaced in April. It answers a contact-centre question in three hedging paragraphs when the desk wanted one line and a referral code. It does both with no source attached, so nobody can tell which edition answered. The 80-page extract of the underwriting handbook pasted into the standing instruction when the assistant was built has never been changed. The handbook runs to 640 pages and is republished quarterly; the wordings and referral schedules add another 1,100 pages. Both moved in April.

A consulting partner has quoted AUD$310,000 and eleven weeks to fine-tune a model on the company’s documents. The chief underwriting officer sponsoring the assistant has that quote, a budget that will not survive being spent twice, and two of her own people arguing that the fix is a document nobody has updated.

What actually matters

Three complaints, three different failures, and sorting them decides the size of the cheque. The flood threshold is a stale fact, wrong now and wrong again after the next quarterly republication. The three paragraphs are a shape problem: nobody told the model the desk wants one line and a referral code. The missing source is a checkability problem that turns the other two into exposure, because an answer without a citation gets acted on, and a delegated authority limit quoted fifteen per cent high produces a claim the company agreed to carry and never priced.

The regulator publishes a bulletin, the flood threshold moves, underwriting republishes the schedule that afternoon, and the question is when the assistant starts saying the new number. Editing an instruction is same-day work. Handing the model the current schedule at question time waits on republishing the document and syncing the index, which is the same working day. Training the number into weights waits for the next training run, which is weeks, and happens again in three months.

The handbook and the wordings come to something like 1.2 million tokens, more than most models on the platform take in a single request, and the 80-page extract travels with every question whether or not it touches those pages. What goes into the window on each request settles which few pages reach the model.

An instruction is owned by a person in underwriting who edits a document. Retrieval is owned by whoever publishes the handbook, plus whoever keeps the index in step with it. A customised model is owned three times over: the labelled training data, the training run, and the serving, and two of those three recur on every base-model upgrade. The sponsor’s real question is the one that decides most AI business cases: who maintains this in year three.

What we’ll filter on

  1. Which failure it fixes. A stale fact or the shape and tone of an answer. These are not interchangeable.
  2. How fast a handbook change reaches an answer. The same afternoon, the same week, or the next training run.
  3. Whether it fits in a single request. Everything sent with a question has to fit the model’s context window alongside the conversation.
  4. Whether the answer can carry a source. A reader who can see the passage and the edition can catch the error the system will eventually make.
  5. Who owns it afterwards. A named person inside the business, with a budget line, or a dependency on the firm that built it.

The landscape

Four instruments sit on a ladder of commitment, not of quality. The full technical ladder has more rungs than this; these are the ones a sponsor signs for.

Rewriting the standing instruction

Prompt engineering at business scale treats the standing instruction as a controlled document: what the assistant is for, who is asking, how long an answer runs, what the desk wants, and what to do when it does not know. Two or three worked examples of the answer shape and an explicit refusal carry most of the value; a model told nothing composes something plausible. Nothing is trained and an edit is live the same afternoon, but every word of it travels with every question, which makes an 80-page extract wasteful as well as wrong. It fixes tone, length, format and refusal, and no facts.

Pasting the whole library into every question

The library is roughly 1.2 million tokens, and whether that fits depends on the model. AWS publishes a 200,000-token context window for Claude Sonnet 4.5 and one million tokens for Llama 4 Maverick, and neither holds it. For Llama 4 Scout it publishes two figures: the Bedrock model card lists ten million tokens, and the launch announcement says Bedrock currently supports three and a half million. Either figure holds the library, and neither Llama 4 model is a choice this insurer can make: AWS has moved both into Bedrock’s Legacy state, and a Legacy model is closed to anyone who has not used it already. Claude Sonnet 5.5 and Claude Opus 5.5, both active, publish a million tokens, which is under the library. Where a window did hold it, the constraint would be the meter rather than the window, because input tokens are metered on every request and this one sends 1,740 pages every time somebody needs one of them.

Retrieval at question time

Retrieval Augmented Generation indexes the handbook and the wordings and looks the answer up at question time: the passages covering flood referral go to the model with the question, so the answer comes from text it has been handed, not from its weights. Amazon Bedrock Knowledge Bases does this as a managed service. It connects to the document stores the company already uses, SharePoint and Amazon S3 among them, re-ingests them on a sync, and includes citations in the answer so the source can be checked. The weights are untouched, so there is no training run, and correction happens by republishing a document. Retrieval does not fix tone, length or format, and a wrong passage produces a confidently wrong answer with a citation on it.

Supervised fine-tuning on company material

Fine-tuning continues training a base model on labelled pairs of a question and the answer the company wanted, and it is the strongest instrument for behaviour, format and task shape. Bedrock also offers reinforcement fine-tuning, which scores responses with reward functions instead of labelled pairs, and distillation, which trains a small model on a larger one’s answers. Each runs on a short list of base models, three of them in the case of reinforcement fine-tuning. Continued pre-training on raw industry text is no longer listed among Bedrock’s customization methods, so treat it as closed. Fine-tuning does not take documents, whatever the quote says: somebody with underwriting judgement writes and checks a few thousand pairs. A base-model upgrade repeats the dataset work and the training run, and none of it makes the April threshold current.

Evaluation

Side by side

Approach Fixes a stale fact Fixes tone and format Correction live the same day Fits in a single request Answer can carry a source Owned inside the business
Rewriting the standing instruction ✗ ✓ ✓ ✓ ✗ ✓
Pasting the whole library into every question ✓ ✗ ✓ ✗ ✓ ✓
Retrieval at question time ✓ ✗ ✓ ✓ ✓ ✓
Supervised fine-tuning ✗ ✓ ✗ ✓ ✗ ✗

No row carries a tick in both of the first two columns, which is why two teams can each be right and still disagree: the flood threshold and the three paragraphs have different answers. The third column settles the facts, since anything that changes quarterly cannot live in a model retrained on a schedule measured in weeks. The fourth removes pasting the library wholesale: the windows big enough for it are on Legacy models the insurer cannot start using, and on those every question would be charged for 1,740 pages of input.

Fine-tuning is not a bad instrument. It addresses one complaint of three, at the highest cost and the slowest correction cycle of the four, and the complaint it addresses is the one a rewritten instruction also fixes the same afternoon.

Which lever fixes which complaint

WHAT CAME BACK WHAT DECIDES IT WHICH LEVER Wrong number Quotes the flood referral threshold that was replaced last April Is it a fact that keeps changing? Retrieval at question time Current the day the schedule is republished and re-indexed Wrong shape Three hedging paragraphs where the desk wanted a referral code Is it tone, length or refusal? The standing instruction Rewritten, versioned, owned, and live the same afternoon No way to check Confident delivery, no source, no record of which edition answered Can the reader see the source? Citation and version stamp Every answer names the passage and the edition it came from The AUD$310,000 quote Fine-tune on company documents, eleven weeks of partner work Narrow, stable, high volume, with a measured baseline? Held back No labelled pairs, no baseline, and the facts still move quarterly One request reads a bounded amount of text. The handbook and the wordings together run to about 1.2 million tokens, more than most models on Bedrock take in one request, so the design question is which few pages travel with the question, not how to send them all.
Three complaints and one quote, each through the gate that decides it: facts to retrieval, shape to the instruction, checkability to citations, and the training spend held back.

The solution

Retrieval for the facts, the standing instruction for the shape of the answer, and the training budget held until there is a narrow task and a measured case for it.

Index the handbook, the wordings and the referral schedules, and hand the assistant the four or five relevant passages with each question instead of the 80-page extract. Underwriting republishes a schedule, the next sync re-indexes it, and the assistant is current without anyone opening a model console. Answers carry the passage and the edition, giving an underwriter a two-second check and compliance something to sample.

Rewrite the standing instruction as a controlled document with a named owner in underwriting, a version number and a change log. It sets the answer shape by channel, one line and a referral code for the contact centre, more room for a technical underwriting question, with two or three worked examples. If the retrieved passages do not cover the question, it says so and names the referral route; a decline is visible and a confident guess is not.

Retrieval moves the failure rather than removing it, so measure at answer level: a monthly sample of real questions checked against the handbook by someone qualified to say. Republication and re-indexing need one owner, because an edition that reaches the intranet on Tuesday and the index the following Monday reproduces the original complaint on a shorter clock.

Fine-tuning keeps a place in the plan, specified well enough for the partner to bid on later. It is worth the money on a narrow task with a stable definition and high volume, where a shorter prompt on a trained model beats a long prompt on a shared one. Routing inbound broker email into one of about forty referral codes changes once a year rather than once a quarter, runs thousands of times a week, and has years of correctly coded email in the mailbox archive as labelled examples. Measure it against what prompting and retrieval already achieve; if it is not materially better, the company saved AUD$310,000 by asking.

Worked example

The regulator publishes at 09:00. By 14:00 the underwriting team has amended the flood referral schedule and republished it, and the owner of republication has started the sync that re-indexes it, because a republished document does not re-index itself. Once that sync finishes, an underwriter asking “what’s the referral threshold for a Zone 3 property” gets the new figure, with the passage and the edition date attached, and nobody touched the assistant. The same change under the vendor’s proposal enters a backlog for the next training run, and until that run completes the assistant answers with the old number.

A week later a handler asks about a commercial property with a partial flood defence that the schedule does not categorise. The retrieved passages are close but not on the question, and the instruction has told the model what to do with that: it says it has no covering rule and names the technical underwriting referral. The handler refers, which the previous assistant would have replaced with a plausible answer nobody could trace.

What’s worth remembering

  1. Three failures, three levers. A stale fact goes to retrieval, answer shape to the standing instruction, checkability to citations; fine-tuning addresses one of the three.
  2. Facts belong in retrieval. Republishing and re-indexing makes the answer current the same day, so anything changing quarterly stays out of a model’s weights.
  3. Prompting fixes shape, not facts. Tone, length, format and refusal change the same afternoon, and every word of the instruction travels with every question.
  4. Send the passages, not the library. 1.2 million tokens exceeds most context windows, so the design question is which few pages travel with each question.
  5. Say when it does not know. The instruction names the referral route for an uncovered question, so a decline is visible rather than a guess.
  6. Fine-tune narrow, stable, high-volume work. Measure what prompting and retrieval already achieve first, so the trained model has a baseline to beat.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.