Exam Room · Advanced Generative AI Developer

Cheat Sheet: Prompt Engineering

· 12 min read

Generative AI Development · part of The Exam Room

A dense revision pass on prompt engineering for Amazon Bedrock: what each technique does, when to reach for it, and which Bedrock feature backs it up.

Techniques at a glance

Technique What it does Use when
Zero-shot Instruction only, no examples The task is common and the output shape is obvious
Few-shot Instruction plus paired example inputs and outputs Output format or edge cases need pinning down
Chain-of-thought Asks for step-by-step reasoning before the answer Multi-step logic, maths, or layered decisions
ReAct Interleaves reasoning with tool actions Tool results have to feed back into the next turn
Tool use (function calling) Response carries a tool call shaped by a schema you supply Your code runs the tool and returns the result
Structured outputs Constrains the response to a JSON schema Downstream code parses the output and cannot retry
System prompt Sets role, tone, and standing context Behaviour must hold across every turn
Delimiters Fence instructions off from input data Untrusted or long user content sits in the prompt
Guardrail input tags Mark which spans a guardrail should evaluate A guardrail must score user text, not your instructions
Prompt templates Parameterised prompt bodies with named variables The same prompt runs with varying inputs
Prompt versions Point-in-time snapshots of a prompt and its config Changes need comparing and switching back
Guardrails Filter prompts and responses against configured policies Safety, privacy, or grounding must be enforced

Decision rules

  • Task is simple and common: zero-shot, no examples.
  • Format keeps drifting: few-shot with a handful of tight examples.
  • Reasoning has several steps: chain-of-thought, and expect more output tokens and latency.
  • Task is trivial: skip chain-of-thought, which adds tokens and latency with no accuracy gain.
  • Tools must run between turns: ReAct. Bedrock does not execute client-side tools; your code runs each call and returns the result.
  • Output must parse every time: structured outputs, either outputConfig.textFormat carrying a JSON schema, or strict: true on the tool definition.
  • JSON asked for in prose fails to parse: move to one of those two, not a firmer instruction.
  • Role, tone, or standing context must persist: the system field, not the first user turn.
  • User content sits beside instructions: delimit it. For Claude models AWS suggests <example> tags around demonstrations.
  • Same prompt runs many times with varying inputs: a Prompt management prompt with variables, invoked by passing the prompt version ARN as modelId.
  • A change needs comparing or switching back: save variants, compare them, then create a version. A version is a snapshot, and you point the application at whichever one you want.
  • Output should repeat closely: lower the temperature. Generation stays stochastic, so identical requests can still differ.
  • Output should vary: raise temperature or Top P, one at a time, so you can tell which moved it.
  • Top K is needed: Converse takes it in additionalModelRequestFields as top_k, not in inferenceConfig.
  • Responses cut off mid-sentence: read stopReason. max_tokens means raise maxTokens; model_context_window_exceeded means the prompt itself is too long.
  • Safety or compliance is required: a guardrail. A system prompt is an instruction, not an enforcement point.
  • Injection gets past your delimiters: guardrail content filters of type PROMPT_ATTACK, plus least privilege on whatever the tools reach.

Traps

  • Treating a system prompt as a security boundary. It shapes output and enforces nothing. Pair it with a guardrail.
  • Assuming plain tool use guarantees a schema-shaped call. Validation comes from strict: true; without it the schema is only a description in the request.
  • Asking for JSON in prose and assuming it parses. stopReason can come back as malformed_model_output.
  • Adding chain-of-thought to trivial tasks. More output tokens, more latency, no accuracy gain.
  • Piling on few-shot examples to close a reasoning gap. Examples fix format, not multi-step logic.
  • Raising temperature to fix wrong answers. Temperature flattens the token distribution and widens the spread; it does not improve accuracy.
  • Confusing Top K and Top P. Top K is a count of candidate tokens, Top P a share of cumulative probability. Lowering either narrows the pool.
  • Forgetting maxTokens caps output only, so truncation reads as a model fault.
  • Running the prompt-attack filter on InvokeModel with no input tags. Untagged content is not evaluated, and that filter needs tags present. Converse uses guardContent blocks instead.
  • Reusing a fixed tag suffix. AWS recommends a fresh random suffix per request, since a static one can be closed off and appended to.
  • Expecting the prompt-attack filter to cover tool traffic. It does not assess toolResult content or the tool definitions in toolConfig.
  • Hardcoding prompts in application code instead of managing and versioning them.
  • Sending inferenceConfig, system or toolConfig alongside a Prompt management prompt in Converse. Those come from the prompt resource instead.
  • Expecting ReAct with no tools defined. Nothing to act on.

Say it in one line

  1. Zero-shot gives instructions only; few-shot adds paired examples to shape output.
  2. Few-shot fixes format and edge cases, not multi-step reasoning.
  3. Chain-of-thought helps layered problems and adds tokens and latency to trivial ones.
  4. ReAct interleaves reasoning and tool calls, and your code runs the tools.
  5. Structured outputs constrains a response to a JSON schema, via outputConfig.textFormat or strict: true on a tool.
  6. JSON requested in prose is unreliable and can fail to parse.
  7. System prompts set role, tone, and standing context across turns.
  8. A system prompt shapes output; a guardrail applies policy.
  9. Guardrails evaluate prompts and responses, and ApplyGuardrail runs them without a model call.
  10. The prompt-attack filter needs input tags to separate user text from your instructions.
  11. Prompt management holds variants for comparison and versions as snapshots you invoke by ARN.
  12. Temperature and Top P shape the sampling pool, Top K caps candidate count, maxTokens caps output length.
  13. Lower temperature repeats more closely, though generation stays stochastic.
  14. Prompt optimization rewrites one short prompt for one target model through OptimizePrompt.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.