The situation
A manufacturer of extruded aluminium window and door systems, 1,100 people across three plants, sells through roughly 900 trade accounts: fabricators, glaziers and builders’ merchants who order through a trade portal. The board has set aside about AUD$400,000 and one seconded manager for AI work this year, and told the operations director to come back with one initiative funded and the rest sequenced or refused.
Four functions have put in an ask.
- Contact centre. Twenty-two agents take about 12,000 calls a month. The head of service has asked why people are calling. Agents pick one of fourteen wrap-up codes at the end of every call, and 38% of calls end up on “Other”. Calls are recorded; none are transcribed. Agents also type free-text notes on roughly half of them.
- Finance. About 5,800 supplier invoices arrive each month from 1,300 suppliers, almost all as PDFs attached to email, in every layout their suppliers happen to use. Three accounts-payable staff key them into the ERP, averaging a little over four minutes each. Around 7% go into a query queue.
- Merchandising. The trade portal carries 4,200 SKUs. The commercial manager wants a “customers who bought this also bought” block on every product page, the way the merchant sites their customers buy from elsewhere do it.
- The line. A plant manager wants surface defects on coated profiles caught before a pallet is wrapped. Scratches, coating pinholes, colour drift. Right now an operator eyeballs profiles at the wrapping station. Around 1 in 60 pallets comes back, and a returned pallet costs roughly AUD$1,800 in freight, rework and credit.
Every one of the four arrived described as “an AI project”, with no capability named and therefore no way to price any of them.
What actually matters
The first thing worth doing is refusing to talk about services. Each of these asks needs a capability, and the capability is what decides the shape of the bill, the shape of the work, and who ends up responsible when the answer is wrong. Named properly, a capability comes with a price shape attached. Document extraction bills per page processed. A foundation model bills per token in and out. A custom model trained and hosted for one factory bills per instance-hour for as long as the endpoint is up, whether or not the line is running. A business intelligence tool bills per seat per month. An operations director cannot compare four asks priced in four different units until somebody has said which unit each one lands in, and that is a naming exercise, not a technical one.
The second is the input. Every capability eats something specific, and it only works if that something already exists in a form a machine can read. Document extraction eats scanned or digital documents, and finance has 5,800 of them a month sitting in a mailbox. Ranking eats interaction history, and the portal has been logging order lines against account numbers for six years. Natural language processing eats text or audio, and the contact centre has both, though the audio is unindexed recordings rather than transcripts. Computer vision eats labelled images, and nobody has ever photographed a defective profile next to a good one and written down which is which. Three of the four asks have their input already. One does not, and the difference between those two states is usually a year.
The third is whether the capability is bought or built, because that decides who owns accuracy after go-live. A packaged capability arrives with the vendor’s accuracy and the vendor’s improvement schedule; the business owns the exception process and nothing else. A built capability means somebody in this building owns a model’s error rate forever, watches it drift as the product mix changes, retrains it, and answers for it when it misses. That is a standing role, not a project cost, and it is the line item that most often goes missing from a business case written by a function that has never run a model.
The fourth is the measure and what a wrong answer costs. Finance already knows minutes per invoice and exceptions per hundred, so a before-and-after number exists whether or not anybody sets out to collect one. The contact centre knows call volume but has no measure of the thing being asked about, since “why people call” has never been recorded in a way anyone trusts. On the line, a wrong answer has direction: a false alarm sends a good pallet for inspection and costs a few minutes, while a miss ships a defect to a customer and costs AUD$1,800. Those two errors are not interchangeable, and any target expressed as a single accuracy percentage hides the miss inside the average.
What we’ll filter on
- Would a query or a rule already answer it? If the answer is a definite value that exists in a system, counting it beats predicting it.
- Does the input exist in a form a machine can read? Not “do we have the data” but “is it captured, joined to something, and machine-readable today”.
- Bought or built? A packaged capability behind an API, or a model this business has to train, host and own.
- Is there a before-and-after number somebody already collects? A measure that predates the project, so nobody has to argue about the baseline afterwards.
- Who owns accuracy once it is live? A named function that watches the error rate, works the exceptions, and decides when to retrain.
The landscape
AI, machine learning and generative AI
Three words get used interchangeably in asks like these, and they mean different things to a budget. AI is the umbrella: any system doing something that would otherwise need human judgement. Machine learning is the subset that gets its behaviour from historical examples rather than from written rules, so it needs a body of past cases with the right answer attached. Generative AI is the subset of machine learning built on foundation models that produce new content, text, images, code, from a prompt, and its distinguishing commercial feature is that the training already happened somewhere else. A generative capability can often be used the afternoon it is switched on, because the model arrives already trained; a conventional machine-learning capability usually cannot, because the training data has to come from this business. That distinction is what separates an ask that starts next month from one that starts after a year of collecting examples.
Ranking and recommendation
Produces an ordered list of items for a specific customer or a specific item. Eats interaction history: who ordered what, when, in what basket. Available packaged, so nobody starts from a blank model. Amazon Personalize ships preconfigured ecommerce recommenders, and two of its use cases are named “Frequently bought together” and “Customers who viewed X also viewed”. A good result is measured on the block itself, as attach rate or basket size on pages that show it against pages that do not. Owned by whoever owns the catalogue, because the failure mode is a commercially silly suggestion rather than a technically wrong one.
Natural language processing over text and speech
Turns unstructured language into something countable: topics, sentiment, entities, categories, summaries. Eats text, or audio that has first been turned into text by transcription, which is its own step and its own bill. Packaged for the common jobs, and a foundation model on Amazon Bedrock covers the less common ones without any training. A good result is measured against a human-labelled sample, because the only way to know whether the categories are right is to have people label a few hundred items and compare. Owned by the function that acts on the output, since a topic nobody acts on is an expensive word cloud.
Computer vision
Classifies or locates things in images and video. Eats labelled images, and the labelling is the project: hundreds to a few thousand examples per defect class, photographed under the lighting the line actually has. General-purpose image services recognise general-purpose things, and none of them has a label for what this factory calls a pinhole, so an industrial-inspection use case is a custom model trained on this plant’s images and hosted on an endpoint that is billed per hour whether the line is running or not. Amazon Lookout for Vision used to be the packaged answer here. It closed to new customers in October 2024 and shut down on 31 October 2025, so it is not selectable. A defect-inspection build now runs through Amazon SageMaker AI or a vision-capable foundation model. A good result is measured as two separate numbers, escapes and false alarms, never one. Owned by quality, with an engineer attached.
Document extraction
Pulls fields, tables and key-value pairs out of documents whose layout varies. Eats PDFs and images, which is what arrives in a finance mailbox all day. Amazon Textract has an AnalyzeExpense operation built for invoices and receipts specifically. Amazon Bedrock Data Automation handles document processing end to end, including classification, extraction, normalisation and validation, with confidence scores attached to the output. Billed per page. A good result is measured in straight-through processing rate and in minutes of keying saved, both of which finance teams already track. Owned by finance, because the exception queue is a finance process and always was.
AI for customer operations
The bundle aimed at contact centres: transcription of calls, summarisation of contacts, categorisation of reasons, assistance to agents while they are on a call. Eats call recordings, chat transcripts and a knowledge base. Whether it arrives packaged depends on the contact platform in use, since these features are increasingly sold as part of the platform rather than as separate services. Measured on handle time, transfer rate and repeat-contact rate. Owned by the service function.
Forecasting and classification
The plain predictive work: a number at a future date, or a label from a fixed set. Eats a clean history where the answer is recorded for past cases, and the shape of the answer does most of the sorting between the two. Built rather than bought, on Amazon SageMaker AI, since these models are specific to one business’s data. Amazon Forecast is closed to new customers, so a demand-forecasting business case that names it is naming something nobody can buy. Measured against a held-out period and against the current method, which is usually a spreadsheet and is often harder to beat than anyone expects.
Reporting
Not an AI capability, and it belongs on the menu because it is the answer to a surprising share of asks. Counting, grouping and charting values a system already holds. Eats structured records. Bought, per seat. On AWS that is Amazon Quick Sight, the business intelligence feature of Amazon Quick, which is what Amazon QuickSight became. A good result is a dashboard somebody opens on a Monday.
Evaluation
Side by side
| Capability | Query or rule would do it | Input exists today | Bought | Baseline number exists | Clear accuracy owner |
|---|---|---|---|---|---|
| Document extraction (invoices) | ✗ | ✓ | ✓ | ✓ | ✓ |
| Ranking and recommendation (portal) | ✗ | ✓ | ✓ | ✗ | ✓ |
| NLP over text and speech (call reasons) | ✗ | ✓ | ✓ | ✗ | ✓ |
| Reporting (call reasons, the coded 62%) | ✓ | ✓ | ✓ | ✓ | ✓ |
| Computer vision (surface defects) | ✗ | ✗ | ✗ | ✓ | ✗ |
Computer vision carries the two crosses that stop a conversation. There is no input, because the images do not exist yet, and no owner, because quality has nobody who has run a model. Reporting carries a tick in the first column, which disqualifies it as an AI initiative and qualifies it as this week’s work for whoever owns the contact centre’s reporting.
The contact centre’s ask splits across two rows, and that split is what makes it fundable. The head of service asked one question, “why are people calling”, and it is two questions. The top reasons by volume are already recorded, because agents pick a wrap-up code on every call, and counting fourteen codes across 12,000 calls is a query. What the codes cannot say is what sits inside the 38% filed as “Other”, roughly 4,500 calls a month, and that slice needs transcription and language processing over recordings and free-text notes. Run as one project, the reporting that a query would produce waits on a model. Split, the reporting side lands in a fortnight and sizes the question that is left.
Naming the capability from how the ask is worded
What the ask sounds like
The words a function reaches for point at the capability it needs, and this is the lookup worth carrying into a budget meeting.
| How a function words it | Capability it needs | What it eats |
|---|---|---|
| “Tell me why customers are contacting us” | Reporting if it is coded, NLP if it is not | Wrap-up codes, or transcripts and notes |
| “Stop us typing these in by hand” | Document extraction | PDFs and scans |
| “Customers who bought this also bought” | Ranking and recommendation | Interaction and order history |
| “Spot the bad ones before they ship” | Computer vision | Labelled images, hundreds to thousands per class |
| “How many will we need next month” | Forecasting | A dated history of the same number |
| “Which bucket does this belong in” | Classification | Past cases with the answer recorded |
| “Answer questions about our own documents” | Generative AI with retrieval | A document corpus worth searching |
| “Have it do the whole task for me” | An agent, with tools and permissions | All of the above, plus a decision on autonomy |
The solution
Fund the invoice work first. Document extraction is packaged, so the timeline is an integration timeline rather than a training timeline; it bills per page, so 5,800 invoices a month prices in an afternoon with the pricing calculator rather than in a workshop; and finance already tracks minutes per invoice and the 7% query rate, so the before-and-after number exists before the project starts and nobody gets to renegotiate the baseline at the end. The failure mode is contained, because a field extracted wrongly lands in an exception queue that three people already staff. Ownership needs no argument either: accuracy on an invoice is a finance number, and the head of AP owns it the way she already owns the query queue.
Confidence thresholds belong in that first business case as a decision finance makes, because the setting decides how many invoices route to a human and how many post unseen. Straight-through processing rate should be the headline measure rather than accuracy, because it converts directly into hours and is hard to argue about.
Merge the contact-centre and merchandising asks into one funded stream behind it. They arrived as two projects from two functions and they run on one body of data: the account, the order lines, the contacts attached to those accounts. Split across two teams, that data gets assembled twice, joined twice, and disagreed about at the first meeting where both sets of numbers are on the same slide. Run as one stream with one owner, the sequence is: ship the call-reason dashboard off the wrap-up codes in a fortnight, which takes a seat licence and somebody’s week; use it to size the 38% that is genuinely unknown; then transcribe and process that slice, and stand up the portal ranking off the same account history. Neither half has a baseline today, so both have to make one: the call-reason dashboard becomes the baseline for the language work behind it, and the ranking block ships to a slice of product pages so attach rate on pages that show it can be read against pages that do not. The commercial manager and the head of service both get their answer, and the business pays for one integration.
Refuse the defect ask as a funded initiative, and give it a camera instead. There are no labelled images, and no amount of money compresses the gap between that and a working inspection model, because the constraint is how long it takes to accumulate enough examples of a pinhole under real line lighting. The honest response is a job rather than a project: mount a camera at the wrapping station, capture every pallet, have the operator who is already inspecting them tag what he finds, and re-price in a year with a labelled set in hand. That is a few thousand dollars and a change to one person’s routine. It also converts the ask into something scoped, because after twelve months the business knows how many defects per class it actually sees, and a business case that says “we have 4,000 labelled examples across three defect classes” prices in a way that “we want AI on the line” never will.
The two capabilities being funded have different afterlives, and that belongs in the allocation rather than in next year’s surprise. Document extraction is bought, so the vendor improves it and the business owns an exception process it already ran. Ranking is bought too, but its output is judged commercially, so somebody has to look at what the portal is suggesting and say whether it makes sense to a glazier. Neither of these needs a data scientist on staff. The defect model, if it is ever built, does, and that hire belongs in its business case rather than in this year’s.
What’s worth remembering
- No price without a named capability. Document extraction bills per page, a foundation model per token, a hosted custom model per instance-hour, reporting per seat.
- Generative AI arrives trained. A conventional machine-learning capability learns from this business’s own history first, and that is most of the timeline gap.
- Count before you predict. A definite value that already exists in a system needs a query or a rule, not a model.
- No machine-readable input, no project yet. It gets a collection job and a re-pricing date, not a budget; the first year is data collection.
- Bought or built decides accuracy. A built capability needs a standing role in its business case, a bought one an exception process.
- Two asks can be one dataset. Differently worded asks over the same data merge into one stream, cheaper than two separate business cases suggest.
Nobody in this scenario opened a console. The operations director walked out of the meeting with a funded initiative, a merged stream, a refused ask that everybody agreed with, and a lookup she can use on the next four. Which service supplies each capability is the next conversation, and it is a shorter one once the capability has a name.