Exam Room · AI Business Strategist

Nobody Here Has Trained a Model

· 31 min read

AI for the Business · part of The Exam Room

The situation

A food-service distributor supplies 11,000 trade customers from four depots: restaurants, pubs, school kitchens, a lot of fish-and-chip shops. Four hundred staff, sixty sales reps on the road with a tablet each, a catalogue of 42,000 lines with a spec sheet and an allergen document behind each, and six years of order history in the ERP. The IT team is fourteen people who look after that ERP and the integrations around it. There are no data scientists or machine learning engineers, and a board paper from the CFO says this programme gets no new headcount.

Two things have already been promised. The sales director told the reps they would get an assistant that answers questions about products and past orders instead of ringing the depot. The finance director told her team they would stop waiting three days for a report; the two-person reporting team runs a queue of about forty requests with a median turnaround of three working days.

The forcing event is neither promise. The contract with the reporting supplier expires in five months, and the renewal on the table is a three-year term at a twenty-two per cent uplift on the AUD$38,000 a year it bills now. Without that date this decision drifts for another year, which is what happened last year.

What actually matters

“How many cases of the five-litre rapeseed oil did Depot 3 sell in week 14” has one right answer sitting in a table, and “is this product gluten free” is a line on a spec sheet that should be shown rather than described. Both are a query rather than a prediction, and a generated sentence in their place is right most of the time instead of every time. What is left is the judged band, smaller than the promises implied: why margin in the north fell last quarter, which oil to offer a customer who complains about price. Those answers get assembled rather than retrieved.

With no ML team, build versus buy is a staffing decision. The routes need different people: somebody to configure a product and own the access rules, application developers, a contract manager, or a team to assemble a training set, fit a model, measure whether it is still accurate as the trade changes, and retrain it when it is not. The fourteen IT people can cover the first three. Nobody can hire the fourth.

Seat pricing is predictable and stops working the moment the audience is not a list of named employees. Consumption pricing has no floor: a six-rep pilot costs pocket money, a rollout costs in proportion to use. Instance pricing bills for the hours the thing exists rather than the hours anyone asks it a question, wasteful at a few hundred questions a day. The reporting audience is about fifteen named people; the assistant is sixty reps now and eleven thousand trade customers once it goes on the website. One programme, two pricing shapes.

Four years of report definitions and saved views sit inside the incumbent’s product, which is why a twenty-two per cent uplift is considered rather than declined. The ERP data, the catalogue and the spec sheets survive any route; the modelling, the prompts, the connectors and the agreed set of correct answers are either held or rented. On Amazon Bedrock the model underneath is a named choice with a published lifecycle. For models launched from 7 September 2026, the model card carries an “EOL no sooner than” date where one applies, and a Legacy period, the notice given once retirement is scheduled: 45 days for some models, six months for most. Migration does not happen automatically. In a packaged product the swap runs on the supplier’s schedule, and the first sign is that answers changed.

What we’ll filter on

  1. Judged, not queried: is the work an assembled answer, or a report and a lookup with a chat window in front of it?
  2. Runs without an ML team: can it be built and kept running by the people already employed, with no new headcount?
  3. First useful answer inside five months, before the reporting contract ends.
  4. Pricing shape matches the audience: seat-based for a bounded named group, consumption for an open one, instance-based only where the volume justifies it.
  5. What the company keeps if the relationship ends: data, definitions, prompts, connectors, the agreed answer set.
  6. What happens when the model underneath is replaced, and whether anyone gets told first.

The landscape

Amazon Quick

A managed assistant over your own data, sold per user, and where AWS points new work now that Amazon Q Business is closed to new customers. Quick Index connects the documents and data sources the business already runs so that answers are grounded in them. Quick Sight, the former QuickSight now folded in, carries the dashboards and answers questions over structured data; Quick Research delivers cited reports; Quick Flows and Quick Automate take on repeated work. An organisation sets it up from the AWS console against IAM Identity Center or IAM Federation. There is no infrastructure to provision, no model to host and no ML expertise required. AWS does not publish which model answers a question; what a user selects is a reasoning mode, Fast, Balanced or Smart, and the heavier modes draw down agent hours faster.

At AWS list prices, seats are USD$20 per user a month on Professional and USD$40 on Enterprise, and both carry a flat USD$250 per account a month on top. Professional pools eight agent hours and 25GB of index storage per user a month and has no overage charge, so the allowance is a ceiling. Enterprise pools eighteen hours and 50GB, then charges USD$3 per agent hour and USD$5 per GB a month above that. A business case built on seats alone leaves out the account fee, and on Enterprise the metered tail with it. The Quick Sight add-ons sit outside the seat price too: in-memory SPICE capacity at USD$0.38 per GB a month with no minimum, formatted multi-page reports at USD$1 per report unit a month with a 500-unit monthly minimum, and alerts and anomaly detection charged on the number of metrics evaluated.

Amazon Bedrock with your own documents behind it

Foundation models from a range of providers through an API, priced by consumption. None of this business’s material sits behind them until Knowledge Bases puts it there: a managed retrieval layer over the catalogue, the spec sheets and the allergen documents. At list that is USD$5 per GB of raw data a month for the index and USD$1 per thousand standard retrieval calls, with the tokens for each answer on top.

Guardrails holds what the assistant must not do: content filters, denied topics, sensitive-information detection, and contextual grounding checks that block an answer the retrieved documents do not support. It is application development rather than data science, so a company with no ML team can own it, and the bill starts near zero and tracks usage.

Amazon SageMaker AI

A model of your own, trained on the company’s own history with the outcome recorded against each case, and then hosted. It suits a narrow repeated prediction whose past decisions are written down. Training bills by the instance-hour, and so does the hosted copy that answers requests, which runs whether or not anybody is asking. A serverless option bills by the millisecond and scales to nothing between requests, which fits a workload that goes quiet. Somebody still has to assemble the training set, fit the model, watch its accuracy and retrain: a standing job this company has been told it may not create.

A vendor product bought through AWS Marketplace

A purchasing channel rather than a fourth technology, buying somebody else’s finished assistant. AWS sells SaaS products as usage-based subscriptions, as upfront commitments, and as free trials, and the charges arrive on the existing AWS bill alongside everything else. A professional services product cannot be bought from its listing at all: the buyer requests a private offer, settles scope and price with the seller, and accepts a contract paid upfront, in instalments, or against payment requests as the work completes. The vendor picks the model, changes it when they choose and owns the roadmap.

Renewing the reporting contract

Signing the renewal is the default. The incumbent answers structured queries and produces the dashboards finance already reads, and the boring baseline wins more often than the people writing strategy papers expect. It does not touch the judged band, has no answer for the sales assistant, and fixes the ownership problem in place for three more years at a price that has just gone up.

Evaluation

Side by side

Option Judged, not queried Runs without an ML team First answer in five months Pricing shape fits the audience Keeps something on exit Model swap comes with notice
Amazon Quick ✓ ✓ ✓ ✓ ✗ ✗
Bedrock with retrieval ✓ ✓ ✓ ✓ ✓ ✓
Custom model on SageMaker AI ✗ ✗ ✗ ✗ ✓ ✓
Vendor product via Marketplace ✓ ✓ ✓ ✓ ✗ ✗
Renew the reporting contract ✗ ✓ ✓ ✓ ✗ ✗

The SageMaker AI row is not a verdict on the service. Neither promise is a narrow repeated prediction with labelled history: nobody has six years of recorded answers to “which oil should I offer this chip shop”. The team does not exist and cannot be assembled and productive inside five months. And a hosted copy serving a few hundred questions a day bills for the hours it sits idle. Move any one of those conditions and the row changes.

Quick and the Marketplace vendor product share a shape: both tick everything up to ownership, because in each case the modelling and the model choice live inside somebody else’s product. That is the right trade for reporting, which the business does not compete on, and the wrong one for product knowledge across 42,000 lines, which it does.

Where each promise lands

THE ASK THE GATES WHERE IT LANDS "Why did depot margin fall in Q3?" 15 named users, structured ERP data "Does it come in 20L, and what did they pay?" 60 reps, then 11,000 "Who will short-ship next month?" six years of labelled history 400 staff, four depots, 42,000 catalogue lines no data scientists, no new headcount five months of reporting contract left Fixed by a query on data we already record? Bounded, licensable audience of named employees? Answer has to appear inside our own application? One narrow prediction, labelled history, funded owner? yes no yes no yes no two of three Report, or a cited lookup Same answer every time, and the spec sheet line shown, not described Amazon Quick Seat pricing, about 15 seats, plus a flat USD$250 account fee each month Bedrock with retrieval Consumption pricing, catalogue behind it, inside the app the reps already use Hold for SageMaker AI History and the prediction are there. Nobody owns it in year two, so it waits
Four gates, and one programme splits into three answers plus a hold. The last gate fails on staffing rather than on data, which is the usual reason.

The solution

Amazon Quick for the finance half, Bedrock with retrieval for the sales assistant, and a written condition under which a custom model becomes the right purchase.

Quick replaces the expiring reporting supplier for a bounded group: nine in finance, four depot managers, the two reporting analysts. Quick Sight carries the dashboards, and the assistant side answers the judged questions over the same modelled data. The budget line is seats plus the account fee plus an estimate of agent hours and index storage, and that estimate picks the plan. Fifteen Professional seats at list are USD$300 a month, USD$550 with the account fee, USD$6,600 a year, and USD$19,800 across the three years the renewal would have run; Enterprise, if the allowance turns out short, is USD$850 a month and USD$30,600 over the same three years. The renewal is AUD$46,360 a year after the uplift, AUD$139,080 over three years. Those are AWS list prices in USD$ against a supplier quote in AUD$, so the figures only line up once converted, and even converted the licence is not what makes this decision. The migration effort and who holds the report definitions afterwards are the larger numbers on the paper.

The migration has to finish before the renewal date, and it fits in five months by moving what is read rather than what exists: count how many reports anybody opens in a quarter, and a few hundred definitions collapse to a dozen. If it slips, take a one-month extension rather than signing three years, and negotiate the end date rather than the unit price.

If any of that dozen is a formatted multi-page report, the add-on carrying that format has a 500-unit monthly minimum: USD$500 a month whatever the volume, taking the Quick bill to USD$1,050 a month and USD$37,800 over three years. The audit records which reports need the format, not only how many reports there are.

The sales assistant goes on Bedrock, with the catalogue, spec sheets and order history behind it in a knowledge base, inside the ordering app the reps already have open. Consumption fits, because there is no seat to sell a customer who logs in twice a year. Two of the fourteen IT people can own the application work, with a fixed-scope engagement alongside them for a quarter, bought through Marketplace onto the existing AWS bill. Write the exit into the statement of work: the prompts, the retrieval configuration, the connectors and the agreed answer set end up the company’s property.

Guardrails carries the assistant’s limits: no allergen or dietary advice generated as prose, because that question routes to the document with the line cited; no pricing or availability commitments, because those come from the ERP or not at all; and a grounding check so an unsupported answer is blocked rather than delivered. AWS lists question answering among the uses that check supports and open-ended chatbot conversation among those it does not, so the rep’s assistant is specified as one question against the retrieved documents at a time.

SageMaker AI is held back under a written condition: one narrow prediction that repeats, a labelled history from a period that still resembles the present, and a named owner with budgeted time in year two. The short-shipment question has the first two, since six years of order lines record which were fulfilled short, and fails the third. Without an owner a model is accurate on day one and degrades unnoticed as the supplier base and product mix move. Revisit at the next budget round with the ownership cost on the paper.

A saved set of about fifty questions with agreed answers covers both halves, re-run whenever a model changes and quarterly otherwise; it turns “the answers feel worse this month” into a number. Both systems go on the approved-tools register with a named owner, because two sanctioned assistants arriving in one quarter is when the unapproved ones multiply in the gaps between them.

Worked example

The finance director asks why margin in the north fell in Q3. Half of that is a query, and Quick answers it from the modelled data the reporting tool used to hold: volume by depot, by category, by month. The other half is assembled, pulling the customer mix and the discount lines into a paragraph she can challenge. Neither half waits three days.

A rep parked outside a chip shop in Kingsway asks whether the five-litre rapeseed oil comes in a twenty-litre drum and what this customer paid last time. The assistant reads the catalogue and the order history and answers with the SKU, the pack size and the last invoice line. The customer then asks whether the batter mix is gluten free. The assistant does not compose an answer; it returns the allergen document with the line highlighted, because a generated sentence about an allergen is the wrong instrument however good the model is.

The depot manager asks which customers are likely to short-ship next month. Nothing answers that today, and it stays on the list until somebody’s job description has it.

What’s worth remembering

  1. Split the queried band out first. A question with one right answer belongs in a report or a cited lookup, which is right every time.
  2. Build, buy or partner is staffing. Ask every route which existing job owns it in year two, and how many days a month that takes.
  3. Audience picks the pricing shape. Seats for a bounded employee group, consumption for an open or external audience, instance-based only where volume fills the hours.
  4. Packaged means somebody else’s model choice. A packaged assistant or vendor product is right where the business does not compete and wrong where it does.
  5. Custom ML: three conditions. One narrow repeated prediction, labelled history, and a funded owner in year two; without the third, SageMaker AI waits.
  6. The model underneath will be replaced. Every route needs the same defence: a saved set of questions with agreed answers, re-run on every change.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.