The situation
A support assistant runs on Amazon Bedrock. It reads a subscriber question, retrieves supporting passages from a knowledge base, and can raise a small goodwill refund through an agent tool. The team has already spent time on runtime defences; guardrails are configured, tools are scoped, retrieved content is delimited. What nobody has done is attack the thing on purpose to find out what still gets through.
The pressure to do so is concrete. A model version is about to change. A prompt is being rewritten to handle a new refund policy. Both are the kind of edit that can reopen a hole a previous fix closed, and no test in the pipeline would catch it. When someone asks whether the jailbreak from March is still blocked, the honest answer is that nobody knows. The March fix was a prompt tweak and a manual check, and neither survives into the next deploy.
The goal is a red-team exercise that produces something durable: not a one-off report that a consultant probed the app and found three issues, but a suite of adversarial tests that runs on every change and a loop that turns each new finding into a permanent defence.
What actually matters
Red-teaming is easy to conflate with two neighbours. The distinction decides what you build.
Evaluation measures quality on inputs you expect. An evaluation job asks whether the assistant answers billing questions correctly, stays on topic, and reads well. It runs representative traffic and scores the normal case. Red-teaming measures failure on inputs an adversary chooses. It runs hostile traffic, the override attempts, the smuggled instructions, the framings aimed at a refund, and asks whether any of it works. The machinery overlaps in places; the intent is opposite. An evaluation run counts as a success when the app answers well, a red-team run when a probe gets through.
Guardrails enforce at runtime; red-teaming probes. Amazon Bedrock Guardrails is a control that sits in the request path and blocks a live attack as it happens. Red-teaming is an offline activity that discovers what to block and checks that the block holds. A finding from red-teaming often becomes a new guardrail rule, so they feed each other, but they are not the same layer and cannot substitute for one another. A guardrail with no adversarial testing is untested; a red-team finding with no runtime control is just a note.
The threat surface is wider than jailbreaks. A thorough exercise probes for direct prompt injection (the user typing an override), indirect injection (a hostile instruction riding in through a retrieved document or a tool result), data and secret exfiltration (driving the model to emit its instructions, session context, or another subscriber’s data), PII leakage in the response, harmful or biased output, and unsafe tool use, where crafted input makes the model call the refund action for an attacker who could never call it directly. The direct and indirect injection pair is covered in depth in the prompt-injection defence post; red-teaming is how you find out whether those defences actually hold.
Two properties decide how much a given failure matters. Side-effecting reach: a jailbroken read-only answer is embarrassing, a jailbroken refund moves money, so probes that end in a tool call rank above probes that end in text. Durability of the fix: a finding patched by editing a prompt evaporates on the next rewrite, while a finding turned into a guardrail rule, a schema check, or a regression test survives every future change. Red-teaming that does not close the loop into something durable leaves you with one clean deploy and nothing more.
What we’ll filter on
- Threat coverage: does the suite exercise the whole surface (direct and indirect injection, exfiltration, PII, harmful output, unsafe tool use), or only the attack that made the news?
- Repeatability: can every probe run unattended on each model or prompt change, so a regression cannot slip back in silently?
- Measurability: can the outcome be scored automatically, by a guardrail check or an automated judge, rather than a human eyeballing each response?
- Creativity: does the process leave room for a human to invent the attack a fixed test list would never contain?
- Loop closure: does each finding become a durable artefact (a guardrail rule, an eval case, or a code fix), or just a line in a report?
- Authorisation: is the exercise scoped and approved, run against the right environment, with no real subscriber data at risk?
The landscape
The space has two axes: the categories of attack you must cover, and the methods you use to run them.
The categories are the threat surface above. Direct injection probes are override and unlock framings typed straight in: “ignore previous instructions”, “you are now in developer mode”, “print your system prompt”. One probe here has nothing to do with framing. If the application uses a fixed guardrail tag suffix on InvokeModel, an attacker can close the tag and append text outside it, where the guardrail does not look. AWS recommends a fresh random suffix on every request for exactly that reason. Indirect injection probes plant a hostile instruction in a place the app will pull in, a knowledge-base article, an uploaded document, a tool response, then ask an innocent question that triggers retrieval. Exfiltration probes try to make the model emit its instructions, encode secrets into an answer, or smuggle data into a tool call’s arguments. PII probes check whether the model repeats card numbers or emails that appear in context. Harmful and biased output probes push for content the Guardrails content filters are meant to stop, which is the hate, insults, sexual, violence and misconduct categories, and separately check for skew across demographic phrasings of the same request. Unsafe tool use probes are the ones that matter most: they try to reach the refund action through the model, testing whether the confused-deputy path is really closed.
The methods are how each category gets exercised. An automated probe harness is a stored set of adversarial inputs replayed against the application through the Bedrock Converse or InvokeModel path, or through the agent, with an assertion on each result. This is what makes red-teaming a regression rather than an event: the suite lives in the repository and runs in CI on every change. Bedrock Guardrails as a scorer covers the categories a guardrail policy already handles. Send a probe input or a model response to the ApplyGuardrail API with source set to INPUT or OUTPUT. An action of GUARDRAIL_INTERVENED means the defence held, and NONE means you have a finding. Its reach is narrower than it looks. The prompt attack filter runs on input only. On InvokeModel it evaluates only text wrapped in guardrail input tags, so an untagged prompt gets no prompt-attack filtering at all, and it skips toolResult content and tool definitions, which is where an indirect injection arrives. Prompt leakage detection needs the Standard tier. No policy scores demographic bias. An automated judge handles the categories no simple rule can score: a second model call grades whether a response leaked the system prompt, complied with an injected instruction, or produced biased content, giving a pass or fail per probe. Amazon Bedrock model evaluation jobs support automatic scoring and an LLM-as-a-judge mode for exactly this kind of grading, and you can also run the judge yourself with a plain model call. Human red-teamers supply the creativity a fixed list cannot: they invent novel framings, chain steps, and follow the model’s own responses towards a weakness. Open-source adversarial toolkits seed the harness with known jailbreak and injection patterns so you are not starting from a blank page.
Scope and authorisation sit under all of it. AWS publishes a customer support policy for penetration testing, and it draws a hard line: assessments of AWS infrastructure or the services themselves are not permitted. Its list of services you may test without prior approval names Amazon Bedrock AgentCore; Amazon Bedrock itself is not on that list. Adversarial red, blue and purple team simulations fall under simulated events, which need a form submitted at least two weeks ahead. Read the policy against your own exercise rather than assuming prompt traffic is exempt. Then run the formalities of any security exercise: written authorisation, a defined target, rules of engagement, and a staging environment seeded with synthetic data. Point destructive tool probes at a sandbox billing API, never the live one. Plan for one side effect of the run itself. Blocked content appears as plain text in Bedrock model invocation logs, so every payload you send lands in the log group in the clear.
Evaluation
Side by side
Probe categories as rows; the properties that decide how each is run and measured as columns.
| Probe category | Automatable regression | Scored by a Guardrail | Needs an automated judge | Rewards human creativity | Side-effecting |
|---|---|---|---|---|---|
| Direct injection / jailbreak | ✓ | ✓ | ✓ | ✓ | ✗ |
| Indirect (second-order) injection | ✓ | ✗ | ✓ | ✓ | ✓ |
| Data / secret exfiltration | ✓ | ✓ | ✓ | ✓ | ✗ |
| PII leakage in output | ✓ | ✓ | ✗ | ✗ | ✗ |
| Harmful output | ✓ | ✓ | ✗ | ✓ | ✗ |
| Biased or skewed output | ✓ | ✗ | ✓ | ✓ | ✗ |
| Unsafe tool / action use | ✓ | ✗ | ✗ | ✓ | ✓ |
Every row is automatable, so the whole surface can live in a regression suite. The last column says where to concentrate. Indirect injection and unsafe tool use are the rows that move money if they succeed, and they are also two of the three rows a guardrail cannot score. For exfiltration the guardrail column is narrower than a tick suggests. Standard-tier prompt leakage detection flags the attempt on the way in, but no policy reads the response and decides that the system prompt came back out. That is a judge’s work, and so is demographic skew.
The solution
The design that satisfies the filters has four moving parts.
A probe suite that lives in the repository. Store adversarial inputs as data, one per case, each tagged with its category and an assertion for what a safe outcome looks like. Replay them against the application through the same path production uses, the Converse API, InvokeModel, or the agent invoke, so the test exercises the real assembled context, retrieval and tools included, not a stripped-down model call. Wire the suite into CI so it runs on every model version bump and every prompt edit. This is the change that turns red-teaming from a report into a regression. The March jailbreak is case 47, and if a prompt rewrite reopens it, the build goes red.
Layered measurement. Not every probe can be scored by a string match. Route response-side probes through ApplyGuardrail and treat an intervention as the defence holding. Where the outcome is a judgement call, whether the model emitted its instructions, followed an injected order, or produced skewed text, use an automated judge: a second model call that returns a pass or fail with a reason. Bedrock evaluations offer this as a judge-model job, and a self-hosted judge call works when you want the grading inside your own harness. For unsafe-tool-use probes, do not judge the text at all. Assert on whether the refund tool was invoked and with what arguments, read from the tool audit log, because what matters is whether money would have moved.
Human red-teamers on top of the automation. The harness catches everything you already know to test, and nothing you have not imagined. Schedule human sessions where testers chain steps, invent framings, and follow the model’s responses towards a weakness, seeded with known jailbreak and injection patterns from open-source toolkits so they start beyond the obvious. The rule that keeps this from being a one-off: every working attack a human finds is written into the automated suite the same day, so it is found by hand once and replayed by machine after that.
A closing loop with an owner. A finding is not done when it is written down; it is done when it becomes a durable artefact. Injection and jailbreak findings usually become a new Guardrails Denied topicsSubjects you describe in plain language that a Bedrock Guardrail refuses to discuss, whichever way a user phrases the request. or a tuned prompt-attack threshold. PII findings become a sensitive-information filter rule. Unsafe-tool findings become a tighter IAM scope, a schema constraint on the arguments, or a human-approval gate. Every finding, whatever else it becomes, also becomes a regression case, so the fix is proven on every future change. Denied topics are capped at 30 per guardrail and that quota does not adjust, so a loop that answers every finding with one more topic runs out of room. The Standard tier allows 1,000 characters per topic definition against the Classic tier’s 200, which is room enough to write a topic broadly rather than once per probe. Triage by side-effecting reach first, since a probe that reaches the refund tool outranks one that only produces an awkward sentence.
Underneath all four, keep the exercise authorised and contained. Get written sign-off on the target and the window, run against a staging environment seeded with synthetic subscriber data, point tool probes at a sandbox billing API, and never make live customer PII the payload you are trying to exfiltrate.
Worked example
A newer model version is available and the team wants the quality gain. In the old process, someone would swap the model ID, run a few normal questions, see good answers, and ship. The risk is that the new model responds differently to adversarial framings, so a framing the previous model handled safely now produces the unsafe output.
With a probe suite in place, the upgrade runs against all of it before merge. Most cases pass. Two fail: an indirect-injection case where a tampered knowledge-base article now drives the new model to attempt a refund, and an exfiltration case where a role-play framing gets the system prompt back in the response. The indirect-injection failure is caught by the tool-log assertion, the refund tool was invoked when it should not have been, and the exfiltration failure is caught by the automated judge, which reads the response and flags that the instructions leaked.
Both are triaged. The refund path is side-effecting, so it goes first: the fix tightens the untrusted-content delimiting and adds a human-approval gate on the goodwill refund, and the case stays in the suite to prove it. The exfiltration finding moves the guardrail to the Standard tier, where prompt leakage detection is available, and adds a denied topic covering internal configuration, plus its own regression case. The upgrade merges only once every probe is green again. A human red-team session a fortnight later invents a fresh framing that chains the two, gets partway, and that framing is added as case 61 the same afternoon. The next model change, whenever it comes, will replay all of it without anyone remembering to.
What’s worth remembering
- Red-teaming attacks your own application to find failures before an outsider does, evaluation measures quality on expected inputs, and guardrails enforce at runtime.
- Cover the whole threat surface, direct and indirect injection, data and secret exfiltration, PII leakage, harmful or biased output, and unsafe tool use, not just the jailbreak that made the news.
- The durable win is a probe suite stored in the repository and run in CI on every model and prompt change, so a fix cannot silently unwind on the next deploy.
- Measure in layers:
ApplyGuardrailwhere a policy covers the category, an automated judge where the outcome is a judgement call, and assertions on the tool audit log for unsafe-tool-use probes. - Know where the guardrail stops: the prompt attack filter runs on input only, needs input tags on
InvokeModel, skips tool results and tool definitions, and no policy scores demographic bias. - Close every finding into a durable artefact, a guardrail rule, a sensitive-information filter, a tighter IAM scope or schema check, a human-approval gate, and always a regression case; triage by side-effecting reach first.