Exam Room · AI Business Strategist

Nine Thousand Refunds a Month

· 29 min read

AI for the Business · part of The Exam Room

The situation

An online homewares and small-electricals retailer despatches around 180,000 orders a month across three countries. About 9,000 of those turn into a refund request that somebody has to decide, and eleven people in the returns team work that queue in the case system. A decision takes six minutes on average: open the case, read the order, read the reason code the customer picked from a dropdown, read what they typed in the free-text box underneath it, check the delivery scan, decide. That is roughly 900 hours a month.

Two proposals arrive in the same budget round. The first writes the published returns policy down as executable rules and lets software approve or refuse without a person. The second trains a model on the four years of closed cases already in the system: about 380,000 decisions, each carrying the order, the reason code, the customer’s free text, the delivery scan, and an outcome that is one of approved in full, partial refund, or refused.

The sponsor is the operations director. She has build budget for one of them and has to choose. Neither slide says what the thing costs in its third year.

What actually matters

The published policy answers a large part of the queue outright: inside the sixty-day window, unopened, tracked as delivered, item value under the threshold, approved; damaged on arrival with a photograph and a delivery scan inside forty-eight hours, approved. A decision fixed by written terms applied to recorded facts is a rule rather than a prediction. The six minutes go on the rest: a kettle that failed at fifty-one days, a box that arrived open with nothing missing. An agent weighs the customer’s history, the product line’s fault rate and the cost of arguing, which is the repeated judgement a model is for.

Refusals are the thinnest of the three outcomes at about six per cent of the four years, roughly 23,000 cases, so even the rare class has examples. A clearance tool approved two months of the pre-Christmas 2024 backlog unread, about 16,000 cases, and those record a queue length rather than a judgement. The March 2026 change moved the returns window from thirty to sixty days, so a model trained across that boundary learns an average of two policies.

Build costs are within a few weeks of each other; ownership is not close. A rule holds until somebody edits it, and the edit is dated and owned. A model’s accuracy changes with nobody touching it, because the model is not wrong, the business moved, and approval rates drift rather than anything breaking. Nobody sees that without outcomes joined back to predictions.

Over-approving is invisible: money leaves, the customer is happy, nothing is logged. Under-approving produces a complaint, a call, and in two of the three countries a right to be told why the claim was refused. A rule cites the clause and the policy version; a model produces a score, and no customer can be told their case fell below a threshold. Where that threshold sits is a business decision rather than a technical one, because the business carries both kinds of error. Nobody measures the agents’ own error rate, so a sample of closed cases and a second reviewer establishes it before anything is built.

What we’ll filter on

  1. Fixed or judged. Does the case have one answer set by written policy on recorded facts, or does it need a judgement no clause enumerates?
  2. Usable history. Are there enough examples of every outcome, recorded as they were decided, from a period that still resembles today?
  3. Three-year cost of ownership. Does the option carry a monitoring and retraining line, and has anybody funded it?
  4. Explainable refusal. Can a customer who is refused be given the reason the decision turned on?
  5. Who catches the error, and when. Does a mistake land in front of a person, or does it show up months later in a cost line?

The landscape

Rules over the published policy

The returns terms become a decision table of conditions, thresholds and outcomes, held where the returns policy owner can edit and version it rather than buried in a code branch. Same input, same output, today and at an audit in two years, and every automated decision stores the clause and policy version it applied. The running cost is the policy owner’s attention when the terms change. A rule cannot answer a case the policy does not cover, and leaving those unanswered is correct behaviour.

A classifier trained on the four years

Amazon SageMaker AI is where a model fitted to your own labelled history gets built, trained and served, here on the 380,000 closed cases minus whatever the three conditions force out. It returns an estimate with a confidence attached rather than a statement, and it is wrong on individual cases and gives no signal about which ones. A confidence score ranks cases against each other; it is not a promise about any one of them, and a high-confidence answer can still be wrong. What serving costs depends on how it is deployed, which at this volume is a live question.

A foundation model reading the free text

Amazon Bedrock gives you a general-purpose model behind an API with no training run and no labelling exercise, priced per token consumed. It handles the language part: summarising the customer’s paragraph, flagging a fault report rather than a change of mind, picking out a mention of a previous return. Your returns policy is not in it unless you put the policy in the request, and the output is still an estimate. Bedrock puts every model in one of three lifecycle states, and the model card carries the notice period before end of life: most models get six months, some get forty-five days. Migration does not happen automatically, so a saved evaluation set is how you tell whether the replacement answers your cases the same way. That move is dated work somebody has to do. A queue worked in office hours does not need an answer in the same second, and AWS prices batch inference on selected models at half the on-demand token rate.

Rules for the clear band, a model on the remainder

Rules answer every case the policy answers, and the residue goes to a model that ranks it and suggests an outcome for an agent to accept or overrule. Certainty where certainty is available, and the model’s mistakes land in front of somebody able to absorb them. It fits a queue holding both kinds of decision, which is most queues.

The eleven agents, unchanged

The baseline every proposal is measured against, and the boring baseline wins more often than a budget round allows for. Costing it honestly means the 900 hours a month and the error rate nobody has measured. Whatever gets built has to beat both.

Evaluation

Side by side

Option Fixed or judged fits Usable history Three-year cost carried Explainable refusal Error caught by a person
Rules over the policy ✓ ✓ ✓ ✓ ✓
Trained classifier alone ✗ ✗ ✗ ✗ ✗
Foundation model alone ✗ ✓ ✗ ✗ ✗
Rules plus model triage ✓ ✓ ✓ ✓ ✓
Agents unchanged ✓ ✓ ✓ ✓ ✓

The two single-model rows fail the first column for the same reason: pointing a model at the band the policy already answers replaces certainty with an estimate and improves nothing. They fail the third column because the monitoring and retraining line is the same size whether the model decides nine thousand cases or eighteen hundred, so aiming it at the whole queue inflates the cost. The classifier alone fails on history where the triage model does not, and the March 2026 window change is the reason: that change decides exactly the cases in the clear band, so across the boundary the history contradicts itself. A model confined to the judgement residue is never asked to learn the window rule. The foundation-model row keeps a tick on history only because it does not use the history at all, which saves the labelling exercise and nothing else.

The agents-unchanged row ticks everything and still loses, on a column that is not in the table: it costs 900 hours a month, and two of the options do the same job for a fraction of that. A row of ticks means an option is admissible rather than best.

Where each case goes

THE QUEUE THE GATES WHERE IT GOES Sealed blender, 22 days, tracked delivered, change of mind Kettle, 51 days, reported faulty after use Box arrived open, nothing missing, not wanted 9,000 refund decisions a month eleven agents, six minutes each Does the published policy give one answer on the recorded facts? Enough history of this case shape, decided the way we decide now? Will the budget carry monitoring and retraining, every year? yes no no yes yes no Rule decides Cites the clause and the policy version. About 7,200 a month Model triages, agent decides Suggested outcome plus the facts behind it. About 1,800 a month Agent decides unaided Same six minutes as today, and no ongoing cost to fund
Three gates split one queue: written policy first, usable history second, and a funded monitoring line third. Failing the third gate is a budget answer, not a technical one.

The solution

Rules for the band the policy answers, a model triaging the remainder, and a named quarterly review that somebody has agreed to pay for.

Replay the four years of closed cases through the draft decision table and count how often the rules reach the outcome an agent did. At eighty per cent the rules carry 7,200 cases a month and the model helps with 1,800; at fifty-five per cent the terms are vaguer than the policy page suggests and rewriting them comes first. The replay takes a week and sizes both proposals.

The decision table belongs to the returns policy owner, not to engineering. She edits the thresholds, the change carries a date and a version, and it reaches production the week it is agreed.

The model ranks and suggests; it does not decide. It returns a proposed outcome with the handful of facts behind it, and the agent accepts in ninety seconds or overrules, so every error lands in front of a person with the case open. The override rate measures the model, with the agent’s decision as the label.

The 16,000 bulk-cleared cases come out of the training data: leaving them in teaches the model to approve whenever a case looks like December. For the March 2026 policy change, either start the training window after it and lose examples, or carry the policy version as an input. Say which you chose in the paper.

At 1,800 predictions a month, worked in office hours, the pricing shape decides the running cost. A SageMaker AI real-time endpoint is a persistent endpoint backed by the instance type you choose, billed by the hour for as long as it is in service, so a queue that goes quiet overnight costs what one running flat out costs. A serverless endpoint bills by the millisecond of compute used to process requests, plus the data processed, and AWS states you pay nothing for idle time. AWS also lists Model Monitor and data capture among the features a serverless endpoint does not support, so the prediction log the quarterly review reads has to be written by the returns application and costed in the build. A serverless endpoint can be converted to a real-time one but not back, which makes it the safer of the two to start on. Reading the free text on Bedrock instead is priced per input and output token on demand, with nothing running between calls.

Fund the monitoring in the same paper as the build: a named owner at roughly a day a month, a dated quarterly review, and three to four days for each retraining cycle. The review reads approval rate by reason code against the same quarter a year earlier, the override rate, and the register of policy and product changes, which triggers the retrain.

Price the three years at one fully loaded agent rate, say AUD$45 an hour. Today’s 900 hours a month is AUD$40,500, about AUD$1.46m over three years. Rules alone leave 180 hours a month, AUD$8,100, about AUD$292,000. Rules plus triage brings the residue to roughly 90 hours, if two cases in three close in ninety seconds and the rest take the full six minutes. That is AUD$4,050 a month. The owner’s day a month plus two retraining cycles adds about AUD$19,000 a year at AUD$1,000 a day, bringing the three years to about AUD$203,000. The two automated options are closer to each other than either is to doing nothing, so choose between them on the ownership line rather than the saving.

If the sponsor will not fund that line, the answer is rules alone and the residual stays with the agents: 7,200 cases a month off an eleven-person team, and about 180 agent-hours against today’s 900. A model nobody has agreed to maintain launches accurate, degrades invisibly, and is trusted the whole time.

Worked example

A sealed blender, returned at twenty-two days, tracked as delivered, reason code “changed my mind”, nothing in the free-text box. Inside the window, unopened, under the value threshold. The rule approves in full, records clause 4.2 and policy version 2026-03, and the case never reaches a queue. If the customer asks why, the clause is on the case.

A kettle bought fifty-one days ago: “it stopped heating last Tuesday”. No rule covers it: the sixty-day window is for change-of-mind returns and this is a fault claim. The model suggests a partial refund and shows the four facts that moved it: fifty-one days elapsed, a fault reason code, two prior refunds from this customer in eighteen months, and a product line whose fault rate sits above its category. An agent agrees and closes it in ninety seconds instead of six minutes.

“Box arrived open, everything’s there but I don’t want it now.” No clause covers it and the history has few close neighbours, so the suggestion comes back with low confidence and the case goes into the ordinary queue unranked, taking the same six minutes as before.

What’s worth remembering

  1. Most queues need both. A decision fixed by written policy on recorded facts is a rule; a repeated judgement no clause enumerates is a model.
  2. Three tests on the history. Enough examples of every outcome, decisions recorded as made rather than cleared in bulk, and a period resembling today.
  3. Estimates need a reviewer. A model is wrong on individual cases without flagging which; skip the person only where mistakes reverse easily.
  4. Ownership is the real difference. Rule and model build costs are comparable; the monitoring and retraining line starts on launch day and never stops.
  5. Drift follows business change. A policy or product change triggers the retrain; the review that catches it needs a named owner and a date.
  6. Pricing shape decides the running cost. At low volume an instance-based endpoint bills idle hours, while consumption-based inference is charged only for what it processes.

Ask who owns the model in year three, what they will look at, and which budget line pays them. A proposal that cannot answer all three is a proposal for rules, whatever the slide says.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.