Exam Room · AI Business Strategist

Pop Quiz: Ninety-Seven Per Cent of What?

· 4 min read

Exam-style

A retailer's fraud screen has run on its web channel for nine months: 1,000,000 transactions, 2% of them fraudulent, 97% accuracy. Audit offers no breakdown and recommends extending it to telephone sales. What should the finance director do with the 97%?

Reveal the answer

C. Send the audit back for the two figures accuracy hides: how many blocked transactions were really fraud, and how many real frauds were blocked

Two per cent of 1,000,000 transactions are fraudulent, so approving everything untouched would be right 980,000 times. That is 98% with no model and no spend, a point above the audit’s 97%. A figure quoted without the rate of the thing being detected mostly reports how rare that thing is. The 99% condition would cost money and catch no more fraud. Say the 97% is 30,000 blocks of which 10,000 were real fraud, with 10,000 frauds waved through. Drop the 20,000 false alarms and the figure is exactly 99%, with the same 10,000 frauds still walking out. A confidence score ranks cases; it is not the probability that a flag is right. What a 0.9 cut-off catches, and what it blocks by mistake, is something to measure. Sample size is not the weakness either: a million transactions is ample. Telephone sales are a different population, so the web figure would not carry across even if it meant something on the web.

AI for the Business · part of The Exam Room

Q. An audit reports 97% accuracy over 1,000,000 transactions, 2% of them fraudulent, and wants the screen extended to a second channel.

A. Send it back. Approving everything untouched scores 98%, so 97% sits below doing nothing. The figure says nothing until it splits into blocks that were really fraud and frauds that got through.

Why? The score is dominated by ordinary purchases correctly left alone, so the headline mostly reports how rare fraud is. Those two errors cost different money. Trading one for the other is a commercial judgement rather than a technical setting. The practitioner version is a classifier that is 99% accurate and blind.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.