A dense revision pass on prompt engineering for Amazon Bedrock: what each technique does, when to reach for it, and which Bedrock feature backs it up.
Techniques at a glance
| Technique | What it does | Use when |
|---|---|---|
| Zero-shot | Instruction only, no examples | The task is common and the output shape is obvious |
| Few-shot | Instruction plus paired example inputs and outputs | Output format or edge cases need pinning down |
| Chain-of-thought | Asks for step-by-step reasoning before the answer | Multi-step logic, maths, or layered decisions |
| ReAct | Interleaves reasoning with tool actions | Tool results have to feed back into the next turn |
| Tool use (function calling) | Response carries a tool call shaped by a schema you supply | Your code runs the tool and returns the result |
| Structured outputs | Constrains the response to a JSON schema | Downstream code parses the output and cannot retry |
| System prompt | Sets role, tone, and standing context | Behaviour must hold across every turn |
| Delimiters | Fence instructions off from input data | Untrusted or long user content sits in the prompt |
| Guardrail input tags | Mark which spans a guardrail should evaluate | A guardrail must score user text, not your instructions |
| Prompt templates | Parameterised prompt bodies with named variables | The same prompt runs with varying inputs |
| Prompt versions | Point-in-time snapshots of a prompt and its config | Changes need comparing and switching back |
| Guardrails | Filter prompts and responses against configured policies | Safety, privacy, or grounding must be enforced |
Decision rules
- Task is simple and common: zero-shot, no examples.
- Format keeps drifting: few-shot with a handful of tight examples.
- Reasoning has several steps: chain-of-thought, and expect more output tokens and latency.
- Task is trivial: skip chain-of-thought, which adds tokens and latency with no accuracy gain.
- Tools must run between turns: ReAct. Bedrock does not execute client-side tools; your code runs each call and returns the result.
- Output must parse every time: structured outputs, either
outputConfig.textFormatcarrying a JSON schema, orstrict: trueon the tool definition. - JSON asked for in prose fails to parse: move to one of those two, not a firmer instruction.
- Role, tone, or standing context must persist: the
systemfield, not the first user turn. - User content sits beside instructions: delimit it. For Claude models AWS suggests
<example>tags around demonstrations. - Same prompt runs many times with varying inputs: a Prompt management prompt with variables, invoked by passing the prompt version ARN as
modelId. - A change needs comparing or switching back: save variants, compare them, then create a version. A version is a snapshot, and you point the application at whichever one you want.
- Output should repeat closely: lower the temperature. Generation stays stochastic, so identical requests can still differ.
- Output should vary: raise temperature or Top P, one at a time, so you can tell which moved it.
- Top K is needed: Converse takes it in
additionalModelRequestFieldsastop_k, not ininferenceConfig. - Responses cut off mid-sentence: read
stopReason.max_tokensmeans raisemaxTokens;model_context_window_exceededmeans the prompt itself is too long. - Safety or compliance is required: a guardrail. A system prompt is an instruction, not an enforcement point.
- Injection gets past your delimiters: guardrail content filters of type
PROMPT_ATTACK, plus least privilege on whatever the tools reach.
Traps
- Treating a system prompt as a security boundary. It shapes output and enforces nothing. Pair it with a guardrail.
- Assuming plain tool use guarantees a schema-shaped call. Validation comes from
strict: true; without it the schema is only a description in the request. - Asking for JSON in prose and assuming it parses.
stopReasoncan come back asmalformed_model_output. - Adding chain-of-thought to trivial tasks. More output tokens, more latency, no accuracy gain.
- Piling on few-shot examples to close a reasoning gap. Examples fix format, not multi-step logic.
- Raising temperature to fix wrong answers. Temperature flattens the token distribution and widens the spread; it does not improve accuracy.
- Confusing Top K and Top P. Top K is a count of candidate tokens, Top P a share of cumulative probability. Lowering either narrows the pool.
- Forgetting
maxTokenscaps output only, so truncation reads as a model fault. - Running the prompt-attack filter on
InvokeModelwith no input tags. Untagged content is not evaluated, and that filter needs tags present. Converse usesguardContentblocks instead. - Reusing a fixed tag suffix. AWS recommends a fresh random suffix per request, since a static one can be closed off and appended to.
- Expecting the prompt-attack filter to cover tool traffic. It does not assess
toolResultcontent or the tool definitions intoolConfig. - Hardcoding prompts in application code instead of managing and versioning them.
- Sending
inferenceConfig,systemortoolConfigalongside a Prompt management prompt in Converse. Those come from the prompt resource instead. - Expecting ReAct with no tools defined. Nothing to act on.
Say it in one line
- Zero-shot gives instructions only; few-shot adds paired examples to shape output.
- Few-shot fixes format and edge cases, not multi-step reasoning.
- Chain-of-thought helps layered problems and adds tokens and latency to trivial ones.
- ReAct interleaves reasoning and tool calls, and your code runs the tools.
- Structured outputs constrains a response to a JSON schema, via
outputConfig.textFormatorstrict: trueon a tool. - JSON requested in prose is unreliable and can fail to parse.
- System prompts set role, tone, and standing context across turns.
- A system prompt shapes output; a guardrail applies policy.
- Guardrails evaluate prompts and responses, and
ApplyGuardrailruns them without a model call. - The prompt-attack filter needs input tags to separate user text from your instructions.
- Prompt management holds variants for comparison and versions as snapshots you invoke by ARN.
- Temperature and Top P shape the sampling pool, Top K caps candidate count,
maxTokenscaps output length. - Lower temperature repeats more closely, though generation stays stochastic.
- Prompt optimization rewrites one short prompt for one target model through
OptimizePrompt.