The situation
An internal productivity assistant for the engineering team needs to do more than answer questions about docs. Common asks:
- “What’s the status of build #4812?” → call the CI API, fetch the build status, summarise.
- “Create a Jira ticket for the bug in the login page.” → call Jira, capture the returned ticket ID, confirm.
- “Page the on-call for the payments team.” → call PagerDuty, trigger the incident.
- “What did we deploy last week?” → call the deploy-history service, filter, summarise.
- “Search our wiki for the runbook on database failover.” → call internal search, retrieve top 3, cite.
Five tools. The assistant should call the right one, pass the correct arguments, confirm before the destructive ones (paging, ticket creation), and hand back structured results the model can weave into a coherent reply.
The team already runs Claude Sonnet 5 on Bedrock. That model has no in-Region on-demand endpoint, so the modelId is a cross-Region inference profile: us.anthropic.claude-sonnet-5 for a US-resident deployment, au. or eu. for the other geographies, global. where residency is unconstrained. They want to ship something this sprint, keep the tooling changeable (add a sixth tool next sprint, tweak a parameter), and avoid building orchestration that duplicates what the API already offers.
What actually matters
Function calling is a protocol between the model and the caller. The caller advertises a set of tools, each with a name, description, and typed arguments. Given a user message, the model returns text, a request to run a tool, or both. On a tool request, the caller executes the function, feeds the result back, and the model either requests more tools or returns a final response.
The first decision is which surface to call through. Some inference APIs offer tool use as a first-class, model-agnostic primitive, one code path regardless of which model is plugged in. Others require model-specific payload shapes. A layer above those, managed agent runtimes take the loop off the caller’s hands entirely and bring a heavier abstraction with them. The right level depends on how many tools there are, how much session state they carry, and how much orchestration the team should own.
The second is how tools are described. JSON schema is the lingua franca: each tool has a name, a description, and typed parameters. Description quality drives tool-choice quality. “Create a Jira ticket.” is less useful than “Create a Jira ticket in the specified project with summary, description, and assignee. Use when the user asks to report a bug, track a task, or file work. Returns the new ticket’s key and URL.”
The third is the execution loop. After the model emits a tool call, the caller parses the call, validates the arguments against the schema, executes the tool, captures the result (or error), packages it back to the model, and re-invokes. The model then either calls another tool or produces a final message. Loop until the model stops calling tools.
The fourth is whether the caller can constrain tool choice: leave it open, require that some tool is called, or name the one tool that must be called. Constraining choice shapes the reply, though conformance to a schema is a separate Bedrock feature rather than a side effect of forcing a tool.
The fifth is confirmation and side-effect safety. Tools with side effects (creating a ticket, paging a human) should surface to the user for confirmation before the tool actually runs. This is pattern-level, not protocol-level: the loop pauses on “destructive” tool calls and waits for user confirmation.
And observability, which sounds optional until the first bad afternoon. Every tool call, name, arguments, result, duration, is a debug signal. When an assistant does the wrong thing, the trace of tool calls explains why.
What we’ll filter on
- Multi-model fit, does this interface work across Claude, Nova, Llama, etc.?
- Schema ergonomics, how easy is it to declare tools and parse calls?
- Orchestration surface, how much loop code we write vs the platform runs?
- Side-effect safety, is there a confirmation gate built in?
- Observability, what traces are emitted without extra code?
The landscape
-
Bedrock Converse API with
toolConfig. The unified interface. PasstoolConfig: { tools: [{ toolSpec: { name, description, inputSchema: { json: {...} } } }, ...] }in the request. When the model requests a tool,stopReasoncomes back astool_useandoutput.message.contentcarriestoolUseblocks withtoolUseId,name, andinput. The caller executes, appends ausermessage holding atoolResultblock with the matchingtoolUseId, and sends the conversation back. Works across Claude, Nova, Llama and Mistral, or any Bedrock model that supports tool use, on the same code path. -
Claude-specific via
InvokeModelwith Anthropic’s Messages payload. Pre-Converse, Claude’s own tool-use schema was reached throughInvokeModelwith a model-specific body. Still works; Converse wraps it. AWS documents one live reason to go direct: the Anthropic-defined tool types (computer_*,bash_*,text_editor_*andmemory_*) are available through the Messages request format rather than throughtoolConfig. -
The AgentCore harness. Declare the model, the instructions, and the tools, and the harness runs the loop, tool execution, memory and tracing, putting the AgentA system that wraps an LLM with tools, memory, and a loop, so it can take multi-step actions toward a goal rather than just answering one prompt. runtime between the caller and the model. Each session runs in an isolated microVM with its own filesystem. Tools reach it through an AgentCore gateway, which publishes Lambda functions and APIs as MCP tools, through a remote MCP server, or as inline functions. It is generally available and it suits long-lived conversations with many tools and heavy session state; for five tools it is a large amount of platform around a loop we can write.
-
LangChain tool-use abstractions. Python-side framework wrapping Bedrock tool use. Cleaner developer ergonomics for some patterns (decorator-style tool definitions, structured parsing). Adds a dependency and an abstraction layer between our code and the API.
-
Custom orchestration around
InvokeModelplain text. Parse a structured response (JSON with tool name and args) from the model’s output. Pre-Converse pattern, brittle, no longer recommended.
Bedrock also has a server-side mode, where it executes the tool itself against a Lambda ARN or an AgentCore Gateway ARN registered as an MCP connector. It runs on the Responses API and, at the time of writing, on the GPT OSS models, so it is not selectable for a Claude assistant and stays out of the comparison below.
Evaluation
Side by side
| Option | Multi-model | Schema ergonomics | Orchestration | Side-effect safety | Observability |
|---|---|---|---|---|---|
| Converse + toolConfig | Native | JSON schema in-line | Caller writes loop | Caller’s job | CloudWatch + CloudTrail |
| InvokeModel (Messages) | Claude only | Anthropic schema | Caller | Caller | Same |
| AgentCore harness | Any harness model | JSON Schema per tool | Managed loop | Inline function gate | Traces built in |
| LangChain tools | Any SDK | Decorator-friendly | Framework | Framework’s | Its own + CloudWatch |
| Custom plain-text | Any | Brittle | Everything | Ours | Ours |
For five tools, a moderate conversational assistant, and a team already using Bedrock Converse elsewhere, the Converse API with toolConfig is the natural fit. It is multi-model, the schema is declarative, the loop is under ~50 lines, and the observability story is the same as any other Bedrock call. A managed agent runtime starts to fit once the tool surface grows to 20 or more tools with heavy session requirements; for five tools it is overkill.
The tool-use loop, in shape
The solution
Tool declarations. Each tool is a JSON spec passed in toolConfig.tools. The description is what tells the model when to use the tool; treat it as PromptThe input you hand to an LLM – system instructions, user message, examples, retrieved documents, tool descriptions, the lot., not documentation. Compare:
Bad: "description": "Create a Jira ticket."
Good: "description": "Create a Jira ticket in the specified project with a summary, description, and assignee. Use when the user asks to file a bug, report an issue, or track a task. Requires the project key (e.g., ENG, SRE) and assignee username. Returns the created ticket's key and URL."
The tool spec also declares inputSchema with typed, constrained parameters. Enums for fixed-value fields (urgency: ["high", "low"]), format strings where sensible (assignee: { type: "string", pattern: "^[a-z]+$" }), required vs optional explicitly marked. Tool calls usually conform, and a mis-typed or missing field turns up often enough that the caller validates before executing. Bedrock’s structured outputs feature adds a strict: true flag on a tool spec that constrains calls to the declared schema, but support is per model and the model card is the register: Claude Sonnet 5 is listed as not supporting structured outputs, so validation in our code stays the backstop.
The loop. In Python:
def run_assistant(user_message, session):
messages = session.get_history() + [{"role": "user", "content": [{"text": user_message}]}]
while True:
resp = bedrock.converse(
modelId=MODEL_ID,
messages=messages,
toolConfig={"tools": TOOL_SPECS},
system=[{"text": SYSTEM_PROMPT}],
)
out_msg = resp["output"]["message"]
messages.append(out_msg)
tool_uses = [c["toolUse"] for c in out_msg["content"] if "toolUse" in c]
if not tool_uses:
return _extract_text(out_msg)
tool_results = []
for call in tool_uses:
if call["name"] in DESTRUCTIVE_TOOLS:
if not session.confirm(call):
tool_results.append({"toolUseId": call["toolUseId"], "content": [{"text": "user declined"}], "status": "error"})
continue
try:
result = dispatch(call["name"], call["input"])
tool_results.append({"toolUseId": call["toolUseId"], "content": [{"json": result}]})
except Exception as e:
tool_results.append({"toolUseId": call["toolUseId"], "content": [{"text": str(e)}], "status": "error"})
messages.append({"role": "user", "content": [{"toolResult": tr} for tr in tool_results]})
50 lines of orchestration; no framework. One response can carry several toolUse blocks (parallel tool use), and the loop covers that by executing all of them and returning the results together.
Destructive-tool confirmation. A runtime set (DESTRUCTIVE_TOOLS = {"create_jira_ticket", "page_on_call"}) intercepts tool calls before execution and asks the user to confirm via the session’s UI (a button in chat, a modal, a Slack approve/deny). Only on approval does the actual call run. On denial, a toolResult saying “user declined” goes back, and the model’s next turn explains the refusal and offers alternatives. The gate is application-layer: Converse has no notion of a destructive tool, so the application marks them.
Error handling. Tool errors (API timeout, 4xx from Jira, invalid arguments that passed schema but failed at execution) go back as a toolResult with status: "error" and text explaining what went wrong. The model’s next turn then retries with different arguments, explains what failed, or escalates. Returning errors as data leaves the conversation open; throwing exceptions up the stack ends it.
Tool choice. toolConfig also accepts toolChoice, a union with three members: auto (the default, tool use optional), any (some tool must be called and no text is generated), and tool with a name (that tool must be called). AWS documents the named form as supported on Anthropic Claude and Amazon Nova models rather than across the board, so a caller written to swap models cannot lean on it.
Observability. Every Converse call publishes CloudWatch metrics under the AWS/Bedrock namespace, dimensioned by ModelId: Invocations, InvocationLatency, InputTokenCount, OutputTokenCount, InvocationClientErrors and InvocationThrottles among them. The response itself carries usage token counts and metrics.latencyMs. Request and response bodies are a separate switch: model invocation logging is disabled by default, and enabling it delivers the full JSON to S3 or CloudWatch Logs. Neither breaks down by tool, so the per-call trace stays an application-layer job, logging the tool name, arguments (redacted where sensitive), result, duration and session ID.
What the request body has to look like
Converse hides the differences between model families, but the request still has a shape that has to be right. Four top-level parts carry the work. messages is an ordered list of turns, each with a role of user or assistant. Each turn’s content is a list of typed blocks rather than a string, and the union is wider than most code uses: text, image, document, video, audio, toolUse, toolResult, guardContent, reasoningContent and cachePoint among them. The system prompt travels in its own top-level field rather than as a first user turn. inferenceConfig holds temperature, topP, maxTokens and stopSequences, and anything outside that set (top_k, for instance) goes in additionalModelRequestFields.
Conversation formatting is where the failures cluster, and they come back as a ValidationException before a token is generated. The common one is a broken tool pairing: a toolResult whose toolUseId matches no toolUse earlier in the conversation. That is why the loop above appends out_msg before it appends anything of its own. toolUseId is capped at 64 characters over a restricted character set, so passing it through as a key elsewhere is fine and truncating it is not. Image and document blocks carry model-specific requirements, so a conversation valid against Claude Sonnet 5 is not automatically valid against whatever gets swapped in next.
stopReason is the other thing worth handling properly. Alongside end_turn and tool_use it can return max_tokens, stop_sequence, guardrail_intervened, content_filtered, malformed_model_output, malformed_tool_use and model_context_window_exceeded. A loop that branches only on the presence of toolUse blocks hands the user a truncated answer without saying why.
InvokeModel has no shared shape at all. The body is the provider’s own JSON schema, so Anthropic’s Messages format, Nova’s, Meta’s and Mistral’s differ in field names and in how tools are declared. A caller written against one of them is a caller written against one model family, which is the argument for routing anything expected to change models without changing code through Converse instead.
SageMaker’s contract is different again. invoke_endpoint takes a serialised payload with an explicit ContentType and Accept, and the container’s own inference handler parses it, so structured data preparation for SageMaker AI endpoints means matching what that handler expects rather than a published API schema. The same fine-tuned model reached through a SageMaker endpoint and through Bedrock takes two different request contracts, and moving between them changes the caller, not just the URL.
What the Lambda behind a tool owes the model
Tool integrations break at the boundary more often than in the model. A standard function definition gets the right tool called with plausible arguments; what happens next is down to the code behind the tool. For these five that code is a Lambda apiece, whether it is dispatched in process or published through a gateway that fronts it as an MCP tool, and each handler owes the model four things.
Validate every parameter against the declared schema on entry rather than assuming the call already conforms. A tool call can carry a date string where an integer was declared, or a null where a required field was, most often after an ambiguous user message. Validation at the top of the handler catches that before anything reaches Jira or PagerDuty.
Return failures as a structured toolResult carrying status: "error" and a short, actionable message: “order id not found, ask the user to confirm it” gives the model its next move, while an exception that escapes the handler kills the turn. The model reads those strings, so they deserve the same care as the tool descriptions.
Make writes idempotent on a caller-supplied key. A retried tool call, whether the retry comes from the model, from the SDK, or from a Lambda invocation that timed out after the write already landed, should not open a second ticket or issue a second refund. Passing the toolUseId through as the idempotency key collapses the retry onto the first result.
Cap the size of what comes back. A search tool that returns fifty full documents might hand back 50,000 tokens. Claude Sonnet 5 has a 1M-token context window, so one such result fits; it is still 50,000 input tokens billed on that call and carried on every turn of the loop after it. Return the top few, trimmed, with an identifier that a follow-up call can expand.
Worked example
User: “The payments team’s on-call, can you page them? We’re seeing 500s on checkout.”
- First Converse call. Response:
toolUse(name=page_on_call, input={team: "payments", urgency: "high", note: "500s on checkout"}). - Caller sees destructive tool. Surfaces confirmation: “Page payments on-call (high) with note ‘Seeing 500s on checkout’?”
- User confirms. Tool executes; PagerDuty returns incident ID
INC-4521. - Caller sends back
toolResult(toolUseId=..., content={incident_id: "INC-4521", url: "..."}). - Second Converse call. Response: text-only. “I’ve paged the payments on-call with a high-urgency incident (INC-4521). The on-call engineer should acknowledge within 5 minutes.”
- Session history now includes the user message, the toolUse, the toolResult, and the final text. Ready for the next turn.
Two Converse calls, one PagerDuty call, one pause for confirmation. Nothing executed that the user had not seen and approved first.
What’s worth remembering
- Use Converse, not InvokeModel, for tool use. Unified across models, clean schema, future-proof.
- Tool descriptions are prompts. Write them like you’re instructing a new team member; which tool gets called follows from what they say.
- The loop is yours but it’s small. ~50 lines covers the common case; frameworks exist but often aren’t needed.
- Destructive tools need a confirmation gate at the application layer. Converse has no notion of a destructive tool, so the application marks and gates them.
- Return errors as data, not exceptions. A structured error leaves the conversation open; a stack trace ends it.
- Handle every
stopReason, not justtool_useandend_turn.max_tokens,content_filteredandmalformed_tool_useall arrive as a normal 200 response.