Exam Room · AI Business Strategist

The Vendor Called It AI

· 27 min read

AI for the Business · part of The Exam Room

The situation

An electrical and plumbing wholesaler runs 22 depots across Australia and holds about 34,000 lines. Roughly 900 times a week a trade counter promises a part that shows as available on the system and is not on the shelf, and the customer buys it two streets away. Finance has costed the leakage at a little over AUD$4 million a year in lost margin. The ask is narrow: by Friday afternoon, a ranked list of which depots will run short of what next week.

Three vendors have answered. The first sells a hosted engine that scores every line at every depot overnight against forty factors, weights configurable by the buyer, priced per depot per month. The second proposes to train a model on the wholesaler’s own three years of despatch history, deploy it into the wholesaler’s AWS account, and hand over a weekly ranked file. The third offers an assistant that reads the depot managers’ weekly notes and the stock report and writes a paragraph per depot on what looks likely to run tight. All three decks carry the word AI on the cover, two in the product name.

The procurement lead has budget to shortlist one and pilot it in a single region of six depots. Nobody in the room can say what makes the three different, other than price.

What actually matters

Five words separate the three decks: algorithm, model, training, inference, prediction. An algorithm is a procedure, a written sequence of steps that turns an input into an output. A model is the artefact you get by running a learning algorithm over data. The first proposal contains an algorithm and no model: forty weights are a procedure somebody wrote down, learned from nothing. The second produces a model. The third calls a model somebody else trained, on text that has nothing to do with these 22 depots.

Those words set the cost shape. Training fits a model to historical examples: it runs offline, consumes compute for the length of the run, and recurs at every refit. Inference is what happens each time the trained model is asked a question, and it runs for the life of the system. Three shapes belong in the business case: consumption-based, charged per token or per request; instance-based, charged per hour of training or serving compute; and seat-based, a fixed fee per licensed unit, here per depot per month. The first year favours the option with no training. The third year is decided by what recurs, the refits and the per-answer charges, and no deck here separates those out.

A rule returns the same outcome every time on the same facts, because nothing is being estimated. A prediction estimates something not yet observed, and it arrives with a number attached: a probability, a score, a confidence. The weekly forecast runs to three-quarters of a million rows and the replenishment team can act on perhaps three hundred, so ranking by that confidence turns the output into a work queue. A sentence saying Bunbury looks tight on 15mm copper carries no such number, so there is no threshold to set on it.

What the buyer supplies decides what remains at the end: forty weights and a service that stops with the subscription, three years of labelled history and an artefact fitted to this business, or documents per request and prose that accumulates nothing.

What we’ll filter on

  1. Where the behaviour comes from: weights a person set, or a relationship learned from the wholesaler’s own data.
  2. What the buyer has to supply, and whether they actually have it: nothing, forty weights, three years of labelled history, or documents at request time.
  3. Cost shape across three years, and whether the three-year total can be quoted before the pilot starts.
  4. Whether the output carries a confidence that a threshold can be set on and a queue ranked by.
  5. What the business owns at the end of the contract, and who is accountable when the answers get worse.

The landscape

Artificial intelligence is the outer set: work normally associated with human judgement, learned or not. Machine learning is the subset whose behaviour comes from data rather than from a programmer, which puts the scoring engine outside it, and generative AI is the part of machine learning that produces content rather than a label or a number. A fuller account of how the terms nest sits behind that.

A hosted scoring engine

Written conditions and weighted factors on a schedule: days of cover, lead time, supplier reliability, a branch manager’s override. Deterministic and traceable, and the boring baseline wins more budget rounds than the vendors pitching against it expect. Nothing in it came from the wholesaler’s history, so it encodes only relationships somebody already knows, and nothing raises an alarm when the trade changes and the weights stay put.

A model somebody else trained

AWS Marketplace lists two kinds of Amazon SageMaker AI product. A model package is pre-trained and needs no further training from the buyer; an algorithm product needs the buyer’s own training data before it predicts anything. Either way the subscription bills as a line item on the monthly AWS bill, not a separate vendor invoice. Negotiated pricing and licence terms come through a private offer, which the seller makes to as many as 25 accounts the buyer names. Under consolidated billing in AWS Organizations, an offer accepted by the management account can be shared with member accounts, though a member account already subscribed has to accept the new offer to get the price. A pre-trained forecasting package pilots in days with no labelling exercise, and holds nothing of this wholesaler’s seasonality, supplier lead times or Kalgoorlie’s three large contractors.

A model trained on the wholesaler’s own history

Amazon SageMaker AI trains, stores and serves a model fitted to the buyer’s own labelled data. A training job is charged for the instance type chosen, for the duration of the run, and the model artefact lands in the wholesaler’s own S3 bucket. Serving is a separate bill, and its shape follows how the model is deployed. A real-time endpoint is charged for the instances hosting it for as long as it exists, busy or idle. Serverless inference is charged for the compute used to process requests, billed by the millisecond, plus the amount of data processed, and it scales to zero between requests. A batch transform job starts instances, writes its predictions to Amazon S3, and stops. Output is a number per line per depot with a confidence, the shape the replenishment queue needs, and labelled history is the entry condition: whether the data is fit to train on decides whether the proposal is buyable at all.

A foundation model reading the notes

Amazon Bedrock is a fully managed service that provides access to foundation models from several providers, with no model to train and no servers to run. As proposed there is no training run and no labelling, and on-demand pricing is charged per input token and per output token. Bedrock does offer customisation, fine-tuning and distillation, either of which would put a training cost back on the buyer; this proposal uses neither. It handles the language well, and it has no access to the despatch history unless the history goes into the request, so what a token costs tracks how much gets sent every week. Its output is content, judged rather than scored.

Evaluation

Side by side

Option Behaviour learned from our data Buyer can supply what it needs Three-year total quotable Output carries a confidence Business owns an asset
Hosted scoring engine ✗ ✓ ✓ ✗ ✗
Marketplace model package ✗ ✓ ✓ ✓ ✗
Model trained on our history ✓ ✗ ✓ ✓ ✓
Foundation-model assistant ✗ ✓ ✗ ✗ ✗

Three of the four rows fail the first column for different reasons: the scoring engine was fitted to nothing, the Marketplace package to somebody else’s data, and the assistant learned language rather than these depots. Pasting last week’s stock report into a request does not change that.

The second column is where the fitted model falls down, and it is the only cell that can be turned around inside a fortnight. Three years of despatch history exists; three years labelled with the shortage being forecast almost certainly does not, because nobody logs the sale they failed to make.

The third column separates the assistant. The other three quote as a fixed annual number that does not move with use: a subscription, a training run plus a scheduled batch, an hourly instance rate. Token consumption moves with what gets sent, and nobody has estimated 22 depots of notes and stock extracts a week.

Reading the three decks

WHAT ARRIVED THE GATES WHAT IT ACTUALLY IS Hosted engine, 40 configurable factors, per depot per month Trained on our three years of despatch history Assistant reads depot notes, writes a forecast paragraph 22 depots, 34,000 lines, 900 lost sales a week one pilot funded, six depots Is the behaviour learned from our own data? Ranked number with a confidence, or written content? Fitted to our despatch history, or to somebody else's? Three years labelled with the shortage we want to predict? no yes content number someone else ours no yes A scoring sheet Deterministic and traceable, no confidence. Keep as the baseline A foundation model Narrates a forecast, does not produce one. Priced per token Somebody else's model Marketplace package. Pilots in days, fitted to another firm's trade Not buyable yet Record the shortage first. No label, no model, at any price Shortlist: fitted to our history, served weekly Training run per refit; batch inference against 22 depots on a Friday Ranked by confidence to the 300 lines the team can act on
Four gates read the three decks: learned or written, number or prose, our history or somebody else's, and whether the label exists at all. The last gate is the one that can stop the purchase.

The solution

Shortlist the proposal that trains on the wholesaler’s own history, contingent on one question answered before anything is signed.

The label is the problem. Despatch history records what left the shelf; the forecast target is what a customer asked for and did not get, and a model fitted to despatch volume predicts what was sold instead. Two proxies are worth checking in the first week: a line-not-picked-in-full flag against picking notes, and a no-sale reason logged at the trade counters for a fortnight. If neither exists, fund the recording before buying any of the three, because a prediction needs something to have been observed.

Quote training and serving separately, because the deck’s single annual figure hides which line grows when the business does. Training is charged for the instances it runs on, for the length of the run, and recurs at the refit cadence. A weekly forecast is a batch job that starts instances on a Friday, writes a file and stops, where a real-time endpoint would be charged for every hour it exists and sit idle for 167 of every 168. Serverless inference, charged by the millisecond of compute plus the data processed, is the line to ask about only if a counter-facing lookup follows later. The public AWS Pricing Calculator turns both lines into an estimate before signing, free, with tax excluded; the in-console one charges USD$2 for a bill estimate after the fifth in a month. Cost Explorer then shows what they actually cost, with up to 13 months of history behind it and a forecast 18 months ahead.

Set the confidence threshold from the operation. The replenishment team can work perhaps three hundred lines a week across 22 depots, nearer eighty in a six-depot pilot, so the threshold is whatever puts eighty rows above the line. The review counts how many of those eighty shortages were prevented, against a baseline recorded before the pilot starts. Six depots are a little over a quarter of the network, so something above AUD$1 million a year of the leakage sits inside the pilot, and that is the figure a three-year quote gets measured against.

Ownership belongs in the contract, not the kickoff: the model artefact, the training code and the feature definitions sit in the wholesaler’s account. Name who is accountable when the answers get worse and fund their time, because a fitted model degrades as lead times, suppliers and the customer mix move away from the years it learned, and the symptom is a slow decline in the hit rate, not an outage.

The other two are ordered rather than rejected. The scoring engine becomes the baseline the model has to beat, and running it in parallel for the pilot quarter costs one subscription. The assistant has a job once there is a ranked list, turning three hundred rows into a paragraph per depot for a Monday morning. Both are second-year conversations.

Worked example

Three lines from the decks, and what each commits the buyer to.

“Our AI engine scores every SKU nightly against forty configurable factors.” An algorithm with no model in it. The forty weights are the product: who owns them, what happens when the trade moves and nobody updates them, and whether anybody notices. The subscription does not move with volume, the easiest of the three to forecast and the hardest to improve.

“We train a proprietary model on your data.” A training run and an artefact. Whose account holds the artefact, how often it is refitted, what a refit costs, and what the outcome column actually is. The last of those can stop the deal, and it is the one most likely to be answered vaguely.

“The assistant reads your depot notes and produces a forecast narrative.” Inference against a model somebody else trained, priced per token in and per token out, returning content rather than a number. If the narrative derives from the notes and the stock report, it summarises documents rather than estimating from three years of outcomes. Ask for the weekly token volume before treating a per-token rate as a budget line.

What’s worth remembering

  1. Algorithm is procedure, model is artefact. A model comes out of running a learning algorithm over data, so a configurable scoring sheet contains no model.
  2. The terms nest. Machine learning is AI whose behaviour is learned from data; generative AI is machine learning that outputs content, not labels or numbers.
  3. What recurs decides the three-year total. Training is compute per run at every refit; inference is charged per answer for the life of the system.
  4. A prediction carries confidence. It estimates something not yet observed, so a queue can be ranked and thresholded; a rule returns an outcome, estimating nothing.
  5. Supply decides ownership. Rules need nothing, a fitted model needs labelled history, a foundation model needs documents each request; only the fitted model remains yours.
  6. No label, no model. A model fitted to your history is unbuyable until the shortage being predicted is recorded; no pricing concession substitutes.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.