Exam Room · AI Business Strategist

Pop Quiz: Same Question, Different Answer

· 4 min read

Exam-style

A facilities company routes 3,800 email enquiries a week to seven teams using a supplier feature built on Amazon Bedrock. One test enquiry landed on Repairs three times in the demo, so the plan promises identical wording always reaches the same team, and checks each release by re-running ten samples. What changes?

Reveal the answer

C. Both commitments. Generation is stochastic, so promise a routing-accuracy rate against a fixed set of labelled cases and put a rules-based check after the model on the categories that have to be certain

Three runs of one enquiry is three samples, not a proof. AWS’s prompt engineering guidance for Bedrock says a response may vary due to the stochastic nature of the generation process, so the same wording can reach two teams with nothing changed in between. A specific instruction with worked examples is the tempting fix; it narrows the spread without closing it, and it does not stop the supplier changing the model or the prompt behind the feature at the next release. Two hundred cases sharpen the estimate, and a sample of any size still measures a rate rather than establishing that identical wording always lands the same way. Score every release against a fixed set of human-labelled expected teams, and keep rules on the categories that must never be wrong.

AI for the Business · part of The Exam Room

Q. A Bedrock routing feature handled one test enquiry correctly three times in a demo. The plan promises identical wording always reaches the same team, and checks each release on ten samples. What changes?

A. Both. Promise a routing-accuracy rate against a fixed, human-labelled set of cases, score every release over that whole set, and put a rules-based check after the model on the categories that have to be certain.

Why? A better prompt narrows the spread without closing it, and the supplier can change the model behind the feature between releases. A bigger sample sharpens the estimate without making the promise true. Where the obligation is to handle something identically every time, a versioned rule gives the same output today and at an audit in two years. Prompt work never will. The practitioner view of the same limit is what it takes to make an LLM output reproducible.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.