Exam Room · Advanced Generative AI Developer

Versioning and Rolling Back Prompts and Models

· 31 min read

Generative AI Development · part of The Exam Room

The situation

A product team runs a customer-facing assistant on Amazon Bedrock. It has a system prompt, a set of few-shot examples, a guardrail that blocks unsafe topics and redacts PII, and it chains retrieval and generation through a Bedrock flow. All of it is wired together in application code: the model id comes from a config entry anyone can edit, the prompt text is a string built in the service, the guardrail is referenced by its working draft, and the flow is invoked through its test alias, which points at the working draft.

Last Tuesday someone widened one few-shot example to cover a new refund case. Quality on unrelated queries dropped a few points over the next two days, and support tickets crept up. Nobody can point at what changed, because three people edited three things that week and none of the edits produced an artefact you can name, diff, or revert. The team’s rollback plan is to remember what the prompt used to say and paste it back.

Worse, the assistant’s tone shifted overnight with no deploy on their side. Someone had moved the config entry to a newer model version after the old one entered its Legacy period, and nothing recorded the change. The underlying problem across all of it is the same: nothing is a version, so nothing is reversible, and no change is deliberate.

What actually matters

Everything here turns on whether a change produces a named, immutable artefact you can point at later. If widening a few-shot example edits a live string, there is no “before” to go back to and no way to prove what the assistant was running on Monday. If it creates a new prompt version, you have a numbered snapshot, the old version still exists, and reverting is selecting the previous number. Every in-place edit has to become a version.

The second concern is drift you didn’t ask for. Bedrock model ids carry the version, and so do cross-Region inference profile ids, so a given id keeps returning the same model rather than sliding to a newer one. Drift comes in through the references around it: a config entry or environment variable that names the model and is edited without a release record, and the lifecycle clock, because a model moves to Legacy and then to end of life, after which requests to it fail. Bedrock publishes an “EOL no sooner than” date and a Legacy notice period on each model card, so a model upgrade can be scheduled rather than discovered.

Third is the blast radius of a change and how fast you can undo it. A change that goes to 100% of traffic the moment it merges reaches everyone before anyone can see a regression, and leaves nothing to switch back to. A change that goes to a small slice first, sits behind a flag, and is gated by an eval-set check and live monitoring, gives you a window to see the regression on real traffic while most users are still on the known-good version. When rollback is repointing an alias at the previous version rather than a redeploy, the window between noticing and recovering is short.

Fourth is coupling between the parts. A generative feature is not one artefact; it is a prompt, a model version, a guardrail version, a few-shot set, and often a flow, and they interact. A prompt tuned against one model version can behave differently against another; a guardrail change can interact with a prompt change. If you version each part independently but ship them separately, you can’t reproduce a known-good combination. The unit that has to be reproducible and revertible is the whole release, the specific combination of versions, not each part in isolation.

And underneath all of it, observability of what is actually running. When a metric moves you want to answer “what version of every artefact was serving this request” without archaeology. That means the running combination is recorded, ideally addressed through one indirection layer (an alias) whose current target you can read at a glance.

What we’ll filter on

  1. Immutability, does the change produce a named version you can diff and revert, or does it edit something in place?
  2. Pinning, is the model id recorded with the release, or resolved at runtime from something anyone can edit?
  3. Rollback speed, is reverting a repointed alias, or a code redeploy and a memory of the old text?
  4. Release coupling, can you reproduce and revert the whole combination of artefacts as one unit?
  5. Rollout control, can the change go to a slice first, gated by evals and monitoring, before it reaches everyone?

The landscape

Model id from mutable config. The model id read from a config entry or environment variable that anyone can change between deploys. The id itself names a version, so the model does not move, but the reference does, and nothing records when. Behaviour then changes with no release to point at.

Pinned model version recorded with the release. A Bedrock model id such as anthropic.claude-sonnet-4-5-20250929-v1:0 already names a specific model version, and a cross-Region Inference profileA Bedrock resource wrapping a model so calls to it can be tagged, routed across regions, or repointed without changing app code. id such as us.anthropic.claude-sonnet-4-5-20250929-v1:0 carries the same version with routing across a geography. Pinning means writing that exact id into the release record rather than resolving it at runtime. Upgrading models becomes a tested change of one identifier, timed against the Legacy and end-of-life dates on the model card.

Prompt as an inline string. The prompt built in application code. Easiest to write, impossible to govern: every edit is silent, there is no version history, and reverting means someone remembering the old wording. This is the state the team is trying to leave.

Prompt management with versions. Store the prompt in Amazon Bedrock Prompt management, with variables for the per-request data, and create versions. Saving gives you a draft version you keep iterating on; a version is a snapshot of the prompt at the moment you created it, numbered from 1. At runtime you pass the prompt version’s ARN as the modelId on Converse or InvokeModel and supply the per-request values in promptVariables, so live traffic runs a fixed version while the draft moves. Reverting a bad wording change is pointing the application at the previous version.

Guardrail draft versus guardrail versions. A Bedrock guardrail has a working draft plus numbered versions, and a version is a snapshot taken while you iterate on that draft. You reference one at invocation with guardrailIdentifier and guardrailVersion, where the version is either DRAFT or a number. Changes to the working draft are not reflected in existing versions, so a deployed version stays put while the draft moves, and rolling back a guardrail regression is naming the prior number.

Flow versions and aliases. A Bedrock flow starts with a working draft (DRAFT) and a test alias (TSTALIASID) that points at it. Versions are numbered from 1 and are immutable, and an alias points at a version. Your application sends InvokeFlow to the alias, so promoting a new version is repointing the alias and rolling back is repointing it at the last-known-good version. The test alias always tracks the mutable draft, which is why production traffic runs on a separate alias.

AgentCore Runtime versions and endpoints. Where the feature is an agent rather than a flow, AgentCore Runtime versions every update automatically, versions are immutable, and endpoints reference a version. The DEFAULT endpoint automatically follows the latest version, so an update moves it with no deploy of yours; a named endpoint has to be updated explicitly, which is what makes a production endpoint reproducible and a rollback one UpdateAgentRuntimeEndpoint call. Amazon Bedrock Agents became Agents Classic and closed to accounts without prior usage on 30 July 2026, so it is not a choice for new work, though its own alias-and-version model behaves the same way for teams already on it.

Staged rollout behind a flag. Put the new release behind a feature flag or a weighted split so it takes a small slice of traffic first, gate promotion on an eval-set check and live production metrics, and keep the previous version one repoint away. This is the operational layer that turns “we created a version” into “we released it deliberately and can pull it back fast”.

One release, all artefacts together. Treat the prompt version, the pinned model id, the guardrail version, and the few-shot set as a single named release recorded together, so you can reproduce and revert the exact combination rather than chasing four independent version numbers.

Evaluation

Side by side

Approach Immutable artefact Pinned to a version Rollback Slice-first rollout Reproduce whole combo
Model id from mutable config ✗ ✗ Edit it back, no record ✗ ✗
Pinned model id in the release ✓ ✓ Change the pinned id ✗ Partial
Inline prompt string ✗ ✗ Remember old text ✗ ✗
Prompt management versions ✓ ✓ Point at prior version ✗ Partial
Guardrail working draft ✗ ✗ Re-edit the draft ✗ ✗
Guardrail versions ✓ ✓ Point at prior version ✗ Partial
Flow versions and aliases ✓ ✓ Repoint the alias ✗ Partial
AgentCore versions and endpoints ✓ ✓ Repoint the endpoint ✗ Partial
Staged rollout behind a flag n/a n/a Flip the flag or repoint ✓ ✗
One release, all artefacts ✓ ✓ Revert the release as a unit ✓ ✓

The bottom row is the target state, and every row above it is a piece of it. Pinning stops the reference moving without a record, versions give you fixed artefacts, the alias or endpoint shortens rollback, the staged flag opens a window to catch a regression early, and bundling the versions into one named release lets you reproduce the exact combination that was serving traffic.

Staged release with alias rollback A new release bundle passes an eval gate, takes a small traffic slice, and an alias points production at the known-good version so rollback is a single repoint. One release bundle Release v7 Prompt version 5 Model id (pinned) Guardrail version 3 Few-shot set B Flow version 12 Eval gate Eval-set check passes? Prod monitoring on the slice Staged rollout 5% slice, flag on healthy for N hrs then widen to 100% prod alias Known-good, one repoint away Release v6 Prompt version 4 Guardrail version 3 Flow version 11 rollback = repoint the alias at v6

The solution

Start with pinning, because it removes the reference nobody is tracking. Write the exact model id, or the exact cross-Region inference profile id where you route across a geography, into the release record instead of reading it from a config entry at runtime. The assistant’s behaviour is then fixed until a release changes the identifier, and a model upgrade becomes a tested change you schedule against the model card’s Legacy and end-of-life dates rather than a tone shift somebody notices on Monday. Convenience is what you are giving up on purpose, so that you can reproduce yesterday.

Move the prompt and its few-shot set into Prompt management and create versions. The prompt gets variables for the per-request data, so the stored artefact is the stable scaffold and promptVariables fills in the input at invoke time. The draft is where you iterate; each version is a numbered snapshot whose ARN live traffic passes as the modelId. Widening a few-shot example is now: edit the draft, create version 5, roll it out, and if quality drops, point back at version 4. The “before” always exists, and the change is diffable rather than a vanished string edit. This is the same approach as treating prompts as tested assets rather than incantations, taken all the way into a managed store with real version numbers.

Do the same for the guardrail. Production should reference a numbered guardrail version, not DRAFT, because edits to the working draft do not reach existing versions and so cannot change what production enforces. Create a version whenever the configuration is one you want to hold still, and a guardrail regression rolls back the way everything else does: name the previous number in guardrailVersion. The draft is for tuning, the version is for serving.

Put the flow behind an alias. Flow versions are immutable; the alias is the indirection your application sends InvokeFlow to. Promoting a new version is repointing the alias, and rollback is repointing it at the last-known-good version, with no code change and no redeploy. Keep the working draft, reached through TSTALIASID, for iteration only, because the draft is mutable and therefore not reproducible. An agent on AgentCore Runtime works the same way with endpoints in place of aliases, with the caveat that the DEFAULT endpoint follows the latest version automatically, so production belongs on a named endpoint you update yourself.

Then wrap the rollout in a gate and a flag. A new release takes a small slice of traffic first, behind a flag or a weighted split, and promotion to full traffic is gated on an eval-set check passing and live production metrics staying healthy on the slice. The previous version stays one repoint away the whole time. This is what turns “we can revert” into “we caught the regression on 5% of traffic and pulled it back in a minute”, because you saw it on real traffic before it reached everyone and the recovery was a single alias change.

Then bundle. Record the prompt version, the pinned model id, the guardrail version, the few-shot set, and the flow version as one named release, so the reproducible and revertible unit is the combination, not four independent numbers. A prompt tuned against one model version can behave differently against another, and a guardrail change can interact with a prompt change, so the thing you promote and the thing you roll back is the whole bundle. When a metric moves, you read one release identifier and know every artefact that was serving the request.

Worked example

Take the change that started the trouble: widening a few-shot example to cover a new refund case. Under the old setup it was an edit to a string in the service, deployed Friday, and by Monday quality on unrelated queries had slipped with no artefact to diff and no “before” to restore.

Replay it as a versioned release. The current bundle is release v6: prompt version 4, the pinned model id, guardrail version 3, few-shot set A, flow version 11, and the production alias points at v6. To make the change you edit the prompt draft, add the refund example to produce few-shot set B, and create prompt version 5. You assemble release v7 from prompt version 5, the same pinned model id, guardrail version 3, few-shot set B, and a new flow version 12. Nothing about v6 has changed; it still exists exactly as it was serving traffic.

You run v7 against the eval set. It passes, so you flip the flag to send 5% of traffic to v7 while 95% stays on v6 through the alias. Production monitoring on the slice is what would have caught Monday’s regression on Friday afternoon: quality on unrelated queries dips on the 5% cohort, well before it reaches everyone. Rollback is repointing the production alias back at v6, and the slice is gone in a minute. Because the whole combination was one named release, you know precisely what moved (prompt version 4 to 5, few-shot set A to B) and precisely what to inspect, rather than three people’s edits across a week with nothing to point at.

Had the eval and the slice stayed healthy, you would widen v7 to 100% by moving the alias, leave v6 in place as the known-good fallback, and the next change would build v8 on top. Every step is deliberate, every step is reversible, and at no point does anyone need to remember what the prompt used to say.

What’s worth remembering

  1. Every edit becomes a version. Without a named artefact there is no “before” to revert to and no proof of what was running.
  2. Pin the model id. An id names one version, so record it with the release and schedule upgrades against the card’s Legacy and end-of-life dates.
  3. Version prompts in Prompt management. The draft is for iteration; live traffic passes a version ARN as the modelId.
  4. Never serve a guardrail from DRAFT. Reference a numbered version, because edits to the draft do not reach versions already deployed.
  5. Rollback is one repoint. Send flows through an alias on an immutable version and agents through a named endpoint, not DEFAULT; no redeploy.
  6. Release the whole bundle. The revertible unit is the combination of prompt, model id, guardrail, few-shot set and flow versions.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.