Exam Room · AI Business Strategist

Cheat Sheet: AI Vocabulary and Solution Types

· 35 min read

AI for the Business · part of The Exam Room

Last-pass revision for AI vocabulary and solution types. Drill the tables, then read the traps twice.

The vocabulary at a glance

Term What it is The business consequence
Algorithm A procedure, written by a person. A configurable scoring sheet is an algorithm with no model in it, however the deck is worded.
Model The artefact produced by running a learning algorithm over data. The asset. Whoever holds it and its training data holds the capability at renewal.
Training Fitting a model to historical examples. One-off compute that recurs at every refit, and impossible where the outcome was never recorded.
Inference One answer for one case. The per-answer cost, and it dominates the three-year total.
Prediction An estimate of something not yet observed. Wrong on individual cases without flagging which. Needs a reviewer or a cheap reversal.
Confidence A score returned alongside the estimate. Makes thresholding possible: doubtful cases to a person, the rest through.
Token A sequence of characters the model treats as one unit of meaning: a whole word, a word part such as -ed, or a punctuation mark. AWS publishes no words-per-token ratio, and each model tokenises differently, so a per-request estimate has to be measured rather than converted from a word count. The billing unit for generative work. Input and output are metered separately, and output costs more per token.
Context window The ceiling on how much text one request can carry. Set by the model rather than the account. More room means changing models, and filling it costs more.

A rule returns an outcome with no confidence attached, because nothing is estimated. That absence separates a rules engine from a model when a vendor will not say which is inside.

AI, ML and GenAI at a glance

“AI” is the umbrella. It covers anything doing work that would otherwise need human judgement, rules included, so a proposal that says only “AI” has not yet said what it will cost.

Family What it covers Buyer supplies Buyer owns afterwards Time to a first useful answer
Machine learning Behaviour learned from examples rather than written by a person. Labelled history from this business, with the outcome recorded on past cases. A fitted model and its data, an asset that survives the supplier. Months. Longer where the labels have to be created.
Generative AI Foundation models that produce content from a prompt. An instruction, and the documents to reason over at request time. Prompts, a retrieval index, an evaluation set. The model is rented and will be replaced. The afternoon it is switched on.

Two asks that sound alike separate on the last two columns, and so do their prices: generative capability arrives already trained somewhere else, so it starts sooner and leaves less behind.

The capability menu at a glance

How the ask gets worded Capability What it consumes Arrives packaged
“Put the right thing in front of them”, “frequently bought together” Ranking and recommendation Interaction history: who ordered what, when, in what basket. Yes.
“Why are people calling us”, “what is in all that free text” Natural language processing Text. Audio is transcribed first, its own step and its own bill. Yes for common jobs. A foundation model covers the rest.
“Spot the defect before it leaves the line” Computer vision Labelled images shot under the site’s actual lighting. No. The labelling exercise is the project.
“Stop keying invoices by hand” Document extraction PDFs and scans of varying layout. Yes, per page, with confidence scores.
“Cut handle time in the contact centre” AI for customer operations Call recordings, chat transcripts, a knowledge base. Increasingly inside the contact platform.
“How much will we sell”, “which of these is fraud” Forecasting and classification History with the answer recorded on past cases. No. Built on this business’s own data.
“I just want to see the numbers” Reporting, not AI Structured records the systems already hold. Yes, per seat. It answers a surprising share of asks.

Input types at a glance

Structured against unstructured describes the shape of the material. Three questions move the price.

The question What sits here What it means for the plan
Can a model read it as it stands? Photographs, call recordings, scanned pages, free-text notes, PDFs. Usable now, with no training and no labelling. Billed per image, page, minute or token.
Must it be pulled into fields first? Line items off an invoice, values off a form, entities out of a note. An extraction step with its own service, its own per-page price and its own confidence scores.
Must its meaning be agreed across systems first? Cause of loss, customer status, “active”, any field two teams have filled in for years. Nothing to build until the definitions are reconciled. This is where a project usually dies.

Two further tests apply to anything predicted rather than read. Did the material exist when the decision has to be made, and was the outcome written down afterwards? A use that fails either is a collection job with an owner and a re-pricing date, and it should be funded as one.

Data quality at a glance

Dimension The specific failure it causes
Completeness Shrinks the usable population. The model learns from whichever cases got filled in.
Accuracy Misleads in a consistent direction, so averaging does not cancel the error.
Consistency of meaning Ranking rows that mean two different things ranks nothing.
Timeliness Answers yesterday’s question, with no signal that the data has aged.
Recorded outcome Nothing to learn from until somebody creates the labels by hand.
Representativeness Fits the model to a population it will not run on, so it is accurate about the wrong world.

Rules or a model at a glance

The case in front of you The deciding condition Answer
Fixed by written policy on recorded facts A clause enumerates it. Rules.
A repeated judgement no clause enumerates The outcome was recorded on past cases as they were decided. A model.
A refusal has to carry its reason Explainability is a stated requirement. Rules, or a model with a person on the refusal.
Being wrong is expensive and nobody notices for months No cheap reversal, no human in the path. Rules, or a model behind a person.
Policy changes several times a year A rule change is a deployment; a model change is a retrain. Rules.
Volume is low and an endpoint would sit idle The running cost follows the pricing shape. Consumption-priced inference rather than an instance-based endpoint.

Most real queues end up using both. Rules take the band the policy settles and a model takes the residue, which costs less than making either one cover the whole queue.

Agent capabilities at a glance

Capability The test
Autonomy Is the next step decided at run time from what the last step returned? If the sequence could have been drawn before the request arrived, it is a workflow.
Tool use Is there a catalogue of calls into live systems, and which of them write? The write list sets the blast radius.
Agent-to-agent communication Does one agent hand a task to another and wait, rather than calling an API? If not, this is one agent however many boxes the diagram has.
Orchestration strategy One agent planning its own work, or a planner routing pieces to specialists? Each component added needs an owner, a budget line and an evaluation set.

Autonomy plus tools changes the approval a system needs. A run assembled at request time is a sequence nobody signed off. The compensating controls belong in the design: a bounded tool catalogue, an approval threshold owned by whoever owns the money, a turn limit, and a queryable trace.

Drift and monitoring at a glance

What moved How it shows up The response
The world: demand, product mix, customer base Accuracy slides with nothing deployed and nothing broken. A scheduled review with a named owner. Retrain on the business change rather than the calendar.
The business: a policy rewritten, a product launched The model answers under rules that were replaced. The change process has to reach whoever owns the model.
The inputs: a renamed field, a stale feed Quality degrades upstream, before the model is involved, and nothing surfaces it. Data checks ahead of the model rather than on its output.
The model underneath: the supplier upgrades the base model Answers move on questions nobody changed. A saved set of questions with agreed answers, re-run on every change.

Tool classification at a glance

Three states, one published set of criteria, and a route between them. Two states fails, because a new tool has nowhere legitimate to sit while somebody decides.

State What it means Owner Review cadence
Approved Cleared for a stated use, with the data classes it may hold named. A named person rather than a committee. An expiry date on every approval.
Under evaluation A legitimate place to sit while a decision is being made. The reviewer, named. A published turnaround, in days.
Blocked Refused, with a named substitute or a date for one. A named person. A re-review date. Without one the block becomes a ban, and the estate nobody can see grows.

The register’s first line has to say what an unlisted tool means, or every reader decides for themselves. Shadow AI is a disclosure problem before it is a security problem, so score the design on whether it makes people likelier to declare what they already run.

Model adaptation at a glance

Three levers, in the order they should be tried.

Lever What it fixes Cost shape Lead time How fast a policy change reaches an answer
Prompting: rewrite the standing instruction Tone, length, format, refusal behaviour. An afternoon to write, then charged as input tokens on every request. Same day. Same afternoon.
Retrieval: hand the current document over at question time A stale or missing fact. An index charged whether or not anyone asks, plus the passages sent on every question. Weeks. On reindex of the republished document.
Fine-tuning: train on company material Weak domain vocabulary, or a narrow, stable, high-volume task with a measured baseline. Training tokens times epochs, plus model storage per month. Serving is per token on demand where the base model and Region allow it, and Provisioned Throughput charged hourly otherwise. Months. The next training run, and the same bill again at every base-model upgrade.

Sort the complaint first. A stale fact, an ugly answer and weak vocabulary are three failures, and only one of them is expensive. Anything that changes quarterly belongs in retrieval rather than in weights.

Prompting well is four moves: say what the output should look like, give one example, name the audience, and state the refusal behaviour. That last one is what makes “say you do not know” beat a confident guess. Keep the standing instruction short, because three extra pages of policy are billed on every future question.

The AWS surface at a glance

Service The condition that selects it Pricing shape
Amazon Bedrock The capability is generative and the model can arrive already trained. Consumption. On-demand per input and output token, batch inference on select models at half the on-demand rate, or Provisioned Throughput charged hourly per model unit on no-commitment, one-month or six-month terms.
Amazon SageMaker AI The prediction is specific to this business’s history and has to be built, trained and hosted. Instance-based. Training per instance-hour, and a real-time endpoint per instance-hour whether or not anything is asking. Serverless inference charges only while a request is processed.
Amazon Quick The answer is a number the systems already hold, charted or asked in plain language. Seat-based, per user per month: by plan across the suite, and by role where only the BI capabilities are wanted. Organisation plans add a flat monthly infrastructure fee per account. Reader access can instead be bought as sessions in bulk, which AWS recommends for an embedded audience or one whose numbers are hard to predict.

Pricing shape moves the total more than unit price does. Seats suit a bounded list of named employees, consumption an open or external audience, and instance hours only a volume that fills them. Custom ML is worth building for one narrow repeated prediction with labelled history and a funded owner in year two, and failing that third condition is a reason to wait.

Standards and frameworks at a glance

Reference The question it answers Certifiable
ISO/IEC 42001 How is AI governed across this organisation, and can that be shown to an outside party? Yes, by an accredited body. AWS holds one covering named services, and your organisation is not certified by association.
ISO/IEC 23053 What do we call the parts of an ML-based system? No. A framework and vocabulary rather than a management system standard.
AWS Cloud Adoption Framework Which organisational capability is missing before this goes past the team that piloted it? No. Six perspectives grouping the capabilities, four transformation domains, four phases.
Well-Architected Responsible AI Lens How was this one AI workload scoped, built, evaluated and operated? No. The lens itself says not to use it as a compliance or assurance checklist.
NIST AI Risk Management Framework What should the risk practice contain at each stage? No. Voluntary, organised as govern, map, measure and manage.

Route by the question on the table rather than by which document somebody has read.

Decision rules

  • Name the capability before pricing the ask. Document extraction is per page, a foundation model per token, a hosted custom model per instance-hour, reporting per seat.
  • If the answer already exists as a value in a system, count it. A query gets there in a fortnight and a model does not.
  • If the outcome being predicted was never written down, there is no model to buy at any price. The honest answer is a collection job with a re-pricing date.
  • Generative capability usually starts the afternoon it is switched on. Conventional machine learning starts with this business’s own history.
  • Stale fact, retrieval. Shape or tone of an answer, the prompt. Weak domain vocabulary on a narrow, stable, high-volume task with a measured baseline, fine-tuning, and only then.
  • Approving an agent means approving the tool catalogue and the worst single call in it. That call is the blast radius the business is accepting.

Traps

  • Accuracy quoted without a base rate. A fraud screen 97% accurate where 2% of transactions are fraudulent is beaten by approving everything, which scores 98%. Ask what the do-nothing baseline scores.
  • A fixed workflow sold as an agent. If the sequence could have been drawn before the request arrived, it is a workflow with a model at one step. Autonomy and tools that write are what change the approval, so agent governance applied to a workflow costs money and removes no risk.
  • A build cost quoted with no monitoring line. Build costs for a rule and a model are often comparable; ownership costs are not. Ask who owns the model in year three, what they look at, and which budget line pays them. A case that cannot answer all three is a case for rules.
  • A demo re-run treated as a test. Generative output is not reproducible by default. One enquiry handled correctly three times is no evidence that identical enquiries get identical handling. An SLA promising identical routing, and a release check that re-runs ten samples both assume otherwise.
  • “We have lots of data” offered as data readiness. Readable as it stands, one agreed meaning where things get compared, present at decision time and outcome recorded are four separate tests, and the field everybody trusts usually fails one.
  • A bigger context window sold as a cost fix. It raises the ceiling without lowering any rate, and filling it costs more.
  • A long input trimmed to fit, with nothing raised. Too much text rarely reaches the business as an error. A layer that trims and answers anyway looks like one that worked. Ask whether long inputs are rejected, trimmed or split, and who decided.
  • A scoring sheet on a deck that says AI. An algorithm is the procedure and a model is the artefact. A configurable weighting may still be the right buy, at a rules-engine price.
  • Reporting funded as an AI initiative. A question answerable by grouping values a system already records is dashboard work. Folding it into a model project makes the cheap half carry the business case for the expensive one.
  • An approval with no expiry. It is a claim about a vendor that stopped being true at some point nobody recorded.
  • Fine-tuning proposed as the first move. It costs an order of magnitude more than retrieval, takes months rather than weeks, and commits the company to owning something new.
  • Drift explained as breakage. Nothing was deployed and nothing failed; the business changed. The retrain trigger is a policy or product change, and the review that catches it needs an owner and a date.
  • ISO/IEC 23053 cited as evidence of governance. It is vocabulary. The certifiable one is 42001, and neither discharges what a regulator requires of a specific use.

Say it in one line

  1. An algorithm is the procedure; a model is the artefact produced by running a learning algorithm over data.
  2. Training is a one-off cost that recurs at every refit; inference is the per-answer cost that dominates the three-year total.
  3. Machine learning needs labelled history from this business; generative AI arrives trained on somebody else’s material.
  4. An AI agent is autonomy over the next step plus tools that change real systems, and its tool catalogue is both the scope document and the blast radius.
  5. Whatever route is chosen, the model underneath will be replaced, so a saved set of questions with agreed answers, re-run on every change, is the defence that survives the supplier.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.