Exam Room · AI Business Strategist

The Assistant That Has to Book the Van

· 27 min read

AI for the Business · part of The Exam Room

The situation

A regional parcel and pallet operator runs 240 vans out of six depots and moves about 38,000 consignments a week. Roughly 1,100 of those fail on the first attempt: nobody in, a locked yard gate, a wrong unit number on an industrial estate.

Since March an assistant on Amazon Bedrock has explained those failures. A customer or a contact-centre agent asks what happened to consignment 4471882, and the assistant reads the consignment record, the scan history and the driver’s free-text note, then writes back in plain language: attempted 11:42, gate locked, no keyholder, returned to the Dandenong depot. About 14,000 conversations a month go through it. It answers well and changes nothing outside the reply: a person reads the answer and acts on it.

Customer operations wants the assistant to rebook the van. Find a slot on a run out of the right depot, move the consignment onto it, confirm to the customer, and where no run reaches the address in time, raise an out-of-pattern job costing the business between AUD$60 and AUD$180. That arrives as a feature on a roadmap and reaches the AI governance group needing a different review from the one the first assistant got.

What actually matters

Four capabilities separate what the operator runs today from what customer operations is asking for, and the approval turns on which of them get switched on rather than on how good the model is. The first is autonomy: today’s assistant runs one fixed path, retrieve the record, write a paragraph, stop. Rebooking branches on whether Monday has capacity, and if not, on whether the customer takes Tuesday, a different depot, or a smaller vehicle. Nobody writes that sequence down in advance, which is the shape of work that gives an agent something to do.

A tool is a call into a live system, and the catalogue is the scope document. Reading the consignment record, a run’s remaining capacity and the depot’s Monday schedule changes nothing in the operator’s systems, so a wrong read produces a wrong answer and no wrong action. Moving a consignment onto a run, confirming by SMS and raising an out-of-pattern job move a vehicle, reach a customer and spend money. The finance director’s question is which of the eleven functions the agent may call, and what the worst single call costs.

Agent-to-agent communication means one agent hands a task to another and gets a result back instead of calling an API: the delivery-scheduling agent asks the customer-contact agent to negotiate a new window. An orchestration strategy answers who plans and who executes, one agent with its own tools or a planner routing pieces to specialists. Both add boundaries between components that each need an owner, a budget, a test suite and a changelog, for work a single agent does here on its own. The operator has three engineers, and each extra component is an annual operating cost rather than a one-off build.

Liability does not move: the operator is liable for a wrongly dispatched van whether a person booked it or a model did. Review moves. Customer operations is asking for 2,700 rebooking conversations a month, and at that volume nobody signs off an individual vehicle decision in advance, so the compensating controls are designed in before it runs.

What we’ll filter on

  1. Blast radius of one wrong action: what does undoing a single mistaken step cost, in money, in a vehicle movement, and in a customer’s day?
  2. Cost per run, and whether it is knowable before the run starts or only after it finishes.
  3. Explainability to an auditor: can somebody reconstruct, six months later, what was decided, by which component, and on what basis?
  4. People to keep it running: how many, doing what, and funded from which line?
  5. Somewhere to put a signature: does the arrangement have a natural place to require human approval on the actions that spend money?

The landscape

The assistant as it stands

Retrieval and an answer, no tools that write. It handles the 11,300 conversations a month that are questions and nothing else, and it stays in service under all three of the others. One model call, cost known in advance, nothing outside the reply changes, and the audit record is the transcript. Every proposal has to beat this baseline.

A fixed workflow with a model at one step

Code owns the sequence: look up the consignment, call the model once to read the driver’s note and classify why the attempt failed, then run the booking steps in a fixed order. The model never chooses what happens next, so every run takes the same path, the cost is written down before it executes, and the audit record is a log of steps that were always going to happen in that order. It handles the roughly 1,900 rebooking requests a month with a standard shape; the rest fall through to a person, which is correct behaviour rather than a gap.

A single agent with a bounded tool set

One agent, a goal in plain language, and a catalogue of eleven functions. It decides which to call and in what order, sees what came back, and decides whether to call another or stop. This is what makes an application agentic, and it handles the 800 rebooking conversations a month that branch. Cost per run varies because the loop runs until the model decides it is done: a request that resolves in three turns and one that takes twelve are the same request to the customer and differ several-fold on the bill. The catalogue bounds the blast radius, not the model’s good behaviour.

A new build here goes to Amazon Bedrock AgentCore. Amazon Bedrock Agents, renamed Bedrock Agents Classic, closed to new customers on 30 July 2026 and now sits in maintenance mode, so it is not selectable for anything starting now.

A planner with specialist agents

A supervisor agent decomposes the request and routes the pieces to specialists: one owning depot capacity, one customer contact, one out-of-pattern jobs and their cost codes. The specialists talk to the supervisor and, in some designs, to each other. The case for it is scale and separation: dozens of tools where one agent’s tool descriptions start colliding, or teams that genuinely own different domains and ship independently. It costs four components rather than one, four sets of prompts and evaluations, and an audit trail in which the component that decided something takes a query to find. At 800 judgement-shaped conversations a month across three engineers, that is a structure built ahead of the problem it solves.

Evaluation

Side by side

Arrangement Blast radius bounded Cost per run knowable Reconstructable for an auditor Runnable by three engineers Natural place for approval
Assistant as it stands ✓ ✓ ✓ ✓ ✓
Fixed workflow, model at one step ✓ ✓ ✓ ✓ ✓
Single agent, bounded tool set ✓ ✗ ✓ ✓ ✓
Planner with specialist agents ✓ ✗ ✗ ✗ ✓

The first two rows tick everything and still lose, on a column the table does not have: neither answers the 800 conversations a month whose steps depend on what the last step returned, and those are what the contact centre spends its day on. A full row of ticks says an arrangement is admissible, not that it is the answer.

The single-agent row loses one tick, on cost, permanently: a loop that runs until the model stops it has no price before it starts, so a turn limit and a monthly ceiling bound it instead. The planner row loses that tick and two more. Reconstructing a decision across four components is a different job from reading one trace, and four is more than three engineers can keep evaluated, versioned and on-call. That is a statement about this operator this year, not a permanent one.

Where each request goes

WHAT COMES IN THE GATES WHAT HANDLES IT Why did my pallet fail, and where is it now? Move Thursday's redelivery to Monday, run 12 has a free slot Rural address, no scheduled run reaches it before Friday 14,000 conversations a month 1,100 failed first attempts a week Does answering it change anything outside the reply? Are the steps fixed before the request arrives? Does one action spend above the AUD$100 threshold? no yes yes no no yes Today's assistant answers Reads the record, writes back. About 11,300 a month Fixed workflow books it Model reads the note at one step. About 1,900 a month Agent acts, trace recorded Eleven functions, nothing else reachable. About 670 a month Agent proposes, dispatcher signs Denied at the gateway until a named person approves. About 130 a month
Three gates split one inbox. The last one is a money question, and it is the only place in the flow where a person has to sign.

The solution

One agent, eleven functions, and an approval step on any spend above AUD$100. It is the third layer rather than a replacement: the existing assistant keeps the 11,300 questions a month, the fixed workflow takes the 1,900 rebookings whose steps are known before the request arrives, and the agent takes the 800 that branch.

The tool catalogue is signed before any code is built: eight reads, and three writes that move a consignment onto a named run, confirm a slot to a customer and raise an out-of-pattern job. Nothing else in the estate is reachable, so the worst a bad run does is move one consignment to the wrong day, tell one customer and spend to the threshold.

Policy in Amazon Bedrock AgentCore evaluates every tool call that goes through the gateway, denying by default and permitting only what a written rule allows. That makes the gateway the enforcement point, so every write the agent can make has to be reached through it for the rule to apply at all. Rules are written in plain English and converted into policy code, validated against the tool catalogue, and can run in logging mode first: evaluated against real traffic and reported, with nothing blocked. They can also read what has already happened in the same conversation, so “an out-of-pattern job above AUD$100 is denied unless a dispatcher approved it first” and a per-conversation spend ceiling are enforced rather than requested, however the model was prompted, which is a limit enforced outside the model.

The threshold belongs to the customer-operations director. At AUD$100, roughly 130 jobs a month need a dispatcher’s signature, about six a working day, in an existing role. At zero, everything queues and the contact centre is where it started. At AUD$150, roughly forty a month need one, about two a working day, and the largest unreviewed job goes up by half.

Cost per run has no ceiling: tokens are billed every turn, and each turn carries the whole conversation. AgentCore bills by consumption with no upfront commitment and no minimum: the runtime for the compute a session uses, the gateway per thousand calls through it, the policy engine per authorisation request at AWS’s published USD$0.000025, and model tokens separately on top. Nothing there carries a seat or a committed spend, so the bill tracks volume rather than headcount. A hard turn limit stops a stuck loop, a monthly budget alerts at seventy per cent, and cost per resolved conversation is reported against the contact-centre minutes it replaced: AUD$1.20 a run against nine dispatcher minutes, so 800 conversations a month is about AUD$960 against 120 dispatcher hours.

AgentCore traces each step, the tool chosen, what went in, what came back and what the agent concluded, and logs policy decisions separately, so a denial shows up as a denial, not an absence. Only the headline metrics arrive by default: step-by-step traces are switched on deliberately, before the run that will be asked about, which makes the audit record something funded at build time rather than found afterwards. The audit answer is one retained query, run in the pilot: given a consignment and a date, every action and the approval behind those above threshold.

Agent-to-agent communication and a planner with specialists stay switched off, and the paper says so, revisited when the catalogue passes about thirty functions or a second team ships on its own schedule.

For the minutes: a few hundred vehicle-scheduling decisions a month are taken by a system whose choices nobody reviewed beforehand, compensated by the catalogue, the threshold, the traces and a weekly review of sampled runs by the operations manager. The governance group approves that sentence, not a productivity improvement.

Worked example

“Redelivery to a farm outside Yea, and the customer needs it before Friday.” No scheduled run reaches that address before the following Tuesday. The agent checks three depots, finds no capacity, checks whether a neighbouring depot’s Thursday run can be extended, finds it cannot, and arrives at an out-of-pattern job quoted at AUD$145.

The gateway denies that call, because no approval for it has been recorded in the conversation. The run pauses and the job goes to the duty dispatcher with the reasoning and the three alternatives that were ruled out. She approves it in forty seconds because the customer is on a contract where a missed window costs the operator more than the AUD$145. The trace records every tool call, both policy decisions, her identity and the timestamp. Nobody wrote that sequence in advance, and everybody can read it afterwards.

The same request with a AUD$60 job attached never pauses. The agent books it inside the catalogue, the turn limit is the only ceiling on it, and the trace is the whole record.

What’s worth remembering

  1. An agent is autonomy plus tools. A model choosing its next step and calling tools that change real systems; the combination changes the approval needed.
  2. The tool catalogue is the scope. The worst single call in it is the blast radius the business is accepting.
  3. Extra agents separate ownership. Agent-to-agent communication and a planner with specialists add components to run and an audit trail that takes a query to read.
  4. Cost per run is not knowable. A turn limit, a monthly ceiling and a cost-per-resolved-conversation figure stand in for a price nobody can quote.
  5. Approval sits outside the agent’s reasoning. A policy boundary enforces it, and the threshold belongs to whoever owns the money rather than the code.
  6. Autonomy means nobody reviews in advance. The governance group should approve that sentence rather than a productivity number.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.