Exam Room · Advanced Generative AI Developer

Handling Ambiguous Questions With Clarification

· 38 min read

Generative AI Development · part of The Exam Room

The situation

A retail company runs a customer-service assistant on Amazon Bedrock, backed by a knowledge base for policies and a set of tools that read the signed-in customer’s account: orders, subscriptions, addresses, payment methods. It handles returns, order status, plan changes, and general policy questions. Most days it works well. The complaints that reach the team are all the same shape.

A customer types “where is my order?” and the assistant picks one of three open orders, usually the wrong one, and reports its status. Another asks “can I cancel?” and gets a full walkthrough of cancelling the whole subscription when they meant a single line item. A third asks “how much will it cost to upgrade?” and gets a number for a plan they are not on. Nothing in the question said which plan they hold, and the assistant answered for one anyway.

None of these are hallucinations in the usual sense. The retrieved policy text is accurate, the tools work, the account data is real. The failure sits upstream of all that. The question was underspecified, and the assistant answered a narrower one that nothing in the input had specified. Nothing in the design checked whether anything was missing. The team wants a system that separates an answerable question from one that only looks answerable.

What actually matters

A language model produces an answer for an ambiguous prompt as readily as for a well-formed one. Given “where is my order” and three candidate orders, the output settles on one reading and says nothing about the other two. The tone is identical whether that reading was right or wrong, so fluency carries no signal about which happened. What the assistant needs is an explicit step that tests sufficiency before it answers, instead of treating every input as answerable.

Concrete signals feed that judgement. Ambiguity is often a matter of missing slots: a return needs an order reference and a reason, a plan change needs which plan and which direction. If a required slot is empty and can’t be inferred, the request is underspecified by construction, and you know that before you call the model. Referential vagueness (“my order”, “the subscription”, “that charge”) is another signal, resolvable only when exactly one candidate exists in the account. And when the answer depends on retrieval, low retrieval confidence is itself evidence. Thin matches, scattered ones, or several documents pulling in different directions all suggest a question that is too broad, or aimed at something the corpus doesn’t cover.

Once ambiguity is detected there are three responses, and choosing between them is the design. The assistant can ask a clarifying question, which is safest when the missing piece can’t be recovered and a wrong answer would do damage. It can offer the most likely interpretations and let the user pick, which is faster than an open question when the candidates are few and enumerable. Or it can resolve the gap from context it already holds: the signed-in customer’s account, the entities named earlier in the conversation. That last one is the best outcome where the context settles the reading, because it asks the user for nothing.

The trade sitting under all of this is over-asking against over-assuming. An assistant that clarifies everything is exhausting and users abandon it. One that assumes everything is confidently wrong and erodes trust faster. The balance moves with how much damage a wrong answer does. Reporting the status of the wrong order wastes a sentence and is easily corrected. Cancelling the wrong subscription or quoting a binding price is not, so those lean towards asking. The blast radius of a mistaken assumption sets how quick the assistant should be to confirm.

Resolution has to be grounded. Filling a missing slot with a plausible value is the original failure in a new place. When the assistant resolves ambiguity, it should do so from real data: an account attribute a tool returned, a document the retrieval step pulled, an entity the user named earlier. If the gap can’t be closed from grounded context, that is the signal to ask.

What we’ll filter on

  1. Ambiguity detection, can the design tell an answerable question from an underspecified one before answering?
  2. Slot and entity completeness, are the required pieces present, and does exactly one candidate resolve a vague reference?
  3. Grounding of the resolution, is a filled gap backed by account data, retrieval, or prior turns rather than a guess?
  4. Damage from a wrong answer, does the response mode scale asking versus assuming to the blast radius of a mistake?
  5. Conversational friction, does it avoid interrogating the user when context already settles the question?
  6. Recoverability, when it does assume, does it state the assumption so a wrong one is easy to correct?

The landscape

Answer directly. When the question is well-specified or context makes the reading unambiguous, just answer. This is the target state for most turns, and everything else exists to reach it safely. The failure is answering directly when the question was not actually clear, which is the situation the team is in now.

Model-judged sufficiency check. Prompt the model to decide whether it has enough to answer before it answers, returning a structured verdict (answerable, or what’s missing) rather than prose. This turns the implicit “just generate something” into an explicit gate, and because it can name the missing piece, it feeds directly into which clarifying question to ask. It adds a reasoning step and is only as good as the prompt, but it catches ambiguity the input-side checks miss.

Custom intent classification. Put a deterministic classifier in front of the model and route on what it returns. Amazon Comprehend handles intent recognition as a custom classification model trained on labelled utterances. ClassifyDocument returns each class with a Score from zero to one, so a low-confidence utterance is routed to a clarifying question instead of to an assistant that will answer anyway. The gate is tuned by a number rather than a prompt, and it classifies without owning the conversation, so it sits in front of a Bedrock flow that already exists. Two constraints come with it. Intents are not among Comprehend’s pre-trained insights, which cover entities, key phrases, PII, dominant language, sentiment, targeted sentiment and syntax, so you train the classifier yourself and keep its labelled set current. Real-time classification also runs against an endpoint you provision in inference units, and the charge for that endpoint continues as long as it is active.

Slot filling. Model the request as a set of required slots and hold the request until they’re filled, prompting for whatever is missing. This is the backbone of conversational designs and is exactly what Amazon Lex does: an intent declares its slots, and the bot elicits any that the utterance didn’t supply before it fulfils the intent. Deterministic and predictable for transactional flows like returns and plan changes; less suited to open-ended questions that don’t decompose into a fixed slot set.

Ask a clarifying question. When something required is missing and can’t be recovered, ask for it in plain language: “Which order do you mean, the trainers or the jacket?” Safest response when a wrong answer does damage, and the most natural when the missing piece is a single fact. Over-used it becomes an interrogation, so reserve it for gaps that context cannot close.

Offer likely interpretations. Rather than an open question, enumerate the candidate readings and let the user choose: “Did you mean cancel the whole subscription, or remove one item from the next box?” Faster than an open prompt when the candidates are few and known, and it doubles as a way to show the user what the assistant can do. It falls apart when the interpretations are many or hard to phrase crisply.

Resolve from account context. Use what you already hold about the signed-in user to settle the ambiguity: if the customer has exactly one open order, “where is my order” has one answer and no question is needed. An agent can call a tool to fetch the missing context, an order list, the current plan, the default address, and resolve the reference from real data. The best outcome when it works, because it’s invisible; the risk is resolving from stale or wrong context, so it needs confirmation when the stakes are high.

Resolve from conversation history. Carry entities named earlier in the session so later vague references bind to them: if the user discussed order #44821 two turns ago, “when will it arrive” refers to that order. Natural in multi-turn chat, and it needs no extra call. The danger is a reference that has drifted, where the user has moved on and the old entity no longer applies, so recency and relevance both matter.

Confident guess with no signalling. Pick a reading and answer as if it were the only one, saying nothing about the assumption. This is the current behaviour and the anti-pattern: it’s indistinguishable from a right answer until the user notices, and it offers no thread to pull to correct it. Even when assuming is the right call, stating the assumption (“showing your most recent order”) turns a silent error into an obvious, correctable one.

Evaluation

Side by side

Approach Detects ambiguity Grounded resolution User friction Best when a wrong answer is Recoverable if wrong
Answer directly ✗ ✗ None Harmless and the reading is clear ✗
Model-judged sufficiency check ✓ Feeds the next step Low Any, as a first gate ✓
Custom intent classification ✓ (low confidence) Routes, doesn’t resolve Low Repeated, phrasable requests ✓
Slot filling ✓ ✓ (declared slots) Medium Transactional flows ✓
Ask a clarifying question ✓ ✓ High Damaging ✓
Offer likely interpretations ✓ ✓ Medium Damaging, few candidates ✓
Resolve from account context ✓ ✓ (account data) None Minor to moderate Partly
Resolve from conversation history ✓ ✓ (prior turns) None Minor Partly
Confident guess, no signalling ✗ ✗ None Never ✗

Read the table against the three complaints. “Where is my order” with several open orders resolves from account context when one order stands out, and by offering the candidates when none does. “Can I cancel” calls for slot filling and a clarifying question, because the blast radius is a cancelled subscription. “How much to upgrade” needs the current plan read from the account before any price is quoted. None of the three is served by the confident guess they currently get.

The decision

Two questions route between the three responses. Can the gap be closed from grounded context, and how much damage does a wrong answer do? The flow below runs on every turn: test sufficiency, try to resolve from context, and only then choose between assuming and asking on the stakes.

Question from the user Enough to answer? Answer directly well-specified, reading is clear Closable from context? Resolve account data, prior turns; state the assumption Wrong answer damaging? Offer readings few, enumerable candidates Ask wrong answer does damage yes no, ambiguous yes no low stakes high stakes

The gates are ordered deliberately. Sufficiency comes first because it is the check the current design skips entirely. Resolution comes before the asking decision, because a gap closed from grounded context asks the user for nothing and should always be tried first. Only when context can’t close the gap does the damage of a wrong answer decide between offering the readings and asking outright.

The solution

Slot filling is the right backbone for the transactional flows, and it is worth using the platform primitive rather than rebuilding it. In Amazon Lex, an intent declares its slots and the bot elicits any the utterance didn’t fill before fulfilment, so “I want to cancel” with no target sits in an elicit-slot state until the customer names what they’re cancelling. That is exactly the gate the “can I cancel” complaint needs. The slot for what to cancel is required, the utterance leaves it empty, and the design does not proceed to a cancellation until it’s filled. Slot filling gives you a deterministic, testable gate for anything that decomposes into required fields, which returns, plan changes, and address updates all do. It fits the structured requests better than the open-ended policy questions, which don’t reduce to a fixed slot set and lean on the model-judged check instead.

An agent reaches the same gate from the other direction, and on AgentCore the exit has to be built rather than switched on. The managed harness takes inline function tools, which execute in your code rather than on the harness, and a clarifying question is one of them. Define ask_subscriber with a single question parameter and write its description as the policy for when to reach for it: whenever a required argument cannot be grounded. When the model calls it, the harness stops and the stream ends with stopReason set to tool_use. Your front end puts the question to the customer, then resumes by invoking the harness again on the same runtimeSessionId. Send back both the assistant’s toolUse message and your toolResult, because the harness does not persist the inline turn, and a stored call with no matching result would leave the session in a corrupted state. Without an exit of this shape the agent has no way to stop and ask, which is much of why a well-instructed agent still emits a plausible order ID it was never given. One thing changes against a platform that detects the missing argument itself. This tool fires only when the model selects it, so how reliably it elicits is prompt work, and it deserves the same evaluation as any other tool-selection behaviour. A separate mechanism runs the opposite way. MCP elicitation lets the tool pause mid-execution and ask the caller for input, and AgentCore Gateway forwards that request to your client. It works only for MCP server targets, and only where the client has declared the elicitation capability.

The model-judged sufficiency check covers what slots can’t. Not every question is a transaction with declared fields; “how does your returns policy work for sale items” is answerable or not depending on whether the corpus covers sale items, and no slot captures that. Here you prompt the model to return a structured verdict, answerable or a named missing piece, before it drafts an answer, and you route on the verdict. Keep the verdict separate from the answer so you can act on it programmatically rather than parsing it out of prose. Retrieval confidence feeds the same gate as a number: the knowledge base’s Retrieve API returns a relevance score on each result alongside the text. “Thin or scattered” then becomes a threshold on the top score and a spread across the rest, both of which you can log, tune against real questions, and test. Below the threshold, narrow the question with the user rather than generate an answer from weak evidence.

Resolving from context suits an agent, provided the resolution comes from a tool result. Given “where is my order”, the assistant calls the order-lookup tool for the signed-in customer and inspects the result. One open order resolves the reference outright and no question is needed; several means the reference is genuinely ambiguous and the design falls through to offering the candidates (“your trainers or your jacket?”). The tool call is what turns a vague pronoun into a grounded entity, and the count of candidates it returns is what decides between resolving silently and asking. The same pattern fixes the pricing complaint: fetch the customer’s current plan before quoting an upgrade, so the number is computed from the plan they actually hold rather than one nothing in the request named. Prior turns feed the same mechanism; an order named earlier in the session binds a later “when will it arrive”, as long as the reference is still recent enough to be the thing the user means.

Not every clarification resolves inside one request. A customer asked which order they meant might answer in two seconds, or close the tab and reply from an email the next morning, and a single synchronous invocation cannot hold a turn open that long. Step Functions is the shape when the exchange has to survive the customer walking away. The state machine runs the sufficiency check, and when the verdict is to ask, a .waitForTaskToken state holds the execution open while the question goes out. It resumes when the reply comes back carrying the task token that state issued. Left alone, a waiting task holds until the execution reaches the one-year quota, so set HeartbeatSeconds or a task timeout and catch the resulting States.Timeout into an escalation branch. An unanswered clarification then falls through to a stated default or reaches a human, instead of sitting open for a year. Keeping the wait, the deadline and the escalation in one state machine puts them where you can inspect them, rather than scattered through retry logic in the front end.

While it is outstanding, the clarification is itself a record: which question went out, which candidates were offered, which slot the answer fills, when it stops being valid. That belongs with the rest of the session, and DynamoDB is the usual home: keyed by session, with a TTL attribute so an abandoned clarification is deleted automatically within a few days of the timestamp you set. Short-term and long-term memory for a chat assistant covers how that store is shaped; a pending clarification is one more item in the short-term half of it.

Whenever you resolve rather than ask, state the assumption. Even a well-grounded resolution can be wrong, a stale default address, an order the customer didn’t mean, so saying it out loud (“showing your most recent order, placed Tuesday”) turns a silent misfire into a one-line correction. It does not reduce the assumption rate. It makes a wrong assumption visible in the turn it happens, which is the difference between an assistant that occasionally guesses and one whose guesses go unnoticed.

Worked example

The customer, signed in, types where is my order? and has three open orders. The current design picks one and reports its status, wrong two times in three.

Under the new flow, the sufficiency step flags the reference as vague: “my order” resolves only if exactly one candidate exists. Before asking anything, the agent calls the order-lookup tool for this customer and reads back three open orders. That count is the routing signal. Three candidates means the reference cannot be resolved silently, so the assistant offers the readings rather than guessing:

You have three orders on the way. Which one do you mean?
  - Trainers, order #44821, out for delivery
  - Jacket, order #44902, in transit
  - Coffee beans, subscription box, ships Friday

If the lookup had returned a single open order, the same flow resolves it silently and answers, stating the assumption so a wrong one is easy to catch: “Your order #44821 (trainers) is out for delivery today.” And if the customer had named an order two turns earlier, the conversation-history binding would have resolved “it” to that order without a tool call at all. One mechanism, the grounded lookup and its candidate count, drives all three outcomes, and none of them is the confident guess the design started with.

What’s worth remembering

  1. Fluency hides ambiguity. A model answers a vague question as confidently as a clear one, so its tone never signals the failure.
  2. Test sufficiency before answering. Add an explicit step that judges whether there is enough to answer, rather than treating every input as answerable.
  3. Three responses to ambiguity. Ask a clarifying question, offer likely interpretations, or resolve from context; choosing between them is the design.
  4. Resolve from grounded context first. Account data or prior turns ask the user for nothing; an agent can fetch what is missing with a tool.
  5. Stakes decide when to confirm. A wrong order status wastes a sentence; a wrong cancellation or price does not, so ask more as damage rises.
  6. State the assumption. When you assume, say so; a silent wrong answer becomes an obvious, correctable one.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.