Exam Room · Advanced Generative AI Developer

Encrypting a Bedrock App End to End With KMS

· 34 min read

Generative AI Development · part of The Exam Room

The situation

A retrieval assistant on Amazon Bedrock is in production. It runs a foundation model behind an API, backed by a Knowledge Base over a vector store, with an agent that calls action-group Lambdas to look up account state and file tickets. The corpus includes contracts and support history; the prompts quote invoice numbers and addresses; the completions get logged for quality review. The team has already been through a security review that covered who can call the model, how the traffic reaches Bedrock, and where the data lives, the ground held by the broader security pass over IAM, PrivateLink, and keys.

What that review deliberately left shallow was the key management. It established that some artefacts should sit under a customer-managed key and moved on. This is the part it moved on from. Compliance now wants a specific thing: for every place customer data or model intellectual property comes to rest, name the encryption key, name who controls its policy, and show that access can be audited and revoked without rebuilding the store. That is a per-artefact exercise, and it comes out differently depending on whether AWS holds the key or you do.

The data is already encrypted. Bedrock encrypts at rest by default and TLS protects everything in transit. The decision in front of the team is narrower. For which artefacts is the AWS-held default enough, and for which do you take the key into your own hands and accept the management work that comes with control, audit and revocation?

What actually matters

The first thing that matters is the difference between the two default key types and a key you create. Specify nothing and Bedrock still encrypts at rest, under an AWS owned key. That key lives in a service-owned account, outside yours. You cannot view it, audit its use in CloudTrail, write a policy for it, or delete it. An AWS managed key sits one step along: created on your behalf, visible in your account, its use logged to CloudTrail, rotated annually, but its policy is set by AWS and you cannot disable it or schedule its deletion. AWS stopped creating that key type for new services in 2021, so Bedrock’s own defaults are AWS owned almost everywhere, and most of the AWS managed keys in this app come from the older supporting services around it. Only a customer-managed key gives you the key policy, the grants, the CloudTrail record of every cryptographic operation, and the ability to disable the key or schedule it for deletion.

The second thing is what a KMS key actually does, because it does not encrypt your gigabytes directly. KMS uses envelope encryption: the service generates a data key, uses that data key to encrypt the bulk data, then asks KMS to encrypt the data key itself under your customer-managed key. The wrapped data key is stored next to the ciphertext; the plaintext data key is used and discarded. To read the data later, the service must call KMS to unwrap the data key, and that call is where your key policy is enforced and where the CloudTrail entry is written. So holding the key does not slow the bulk data down. The KMS calls scale with the number of data keys, not with the gigabytes sitting under them.

The third thing is that the key policy, not an IAM policy alone, is the real gate on encrypted data. A KMS key carries its own resource policy, and for a customer-managed key that policy is the authoritative statement of who may use the key to decrypt. Grants are the fine-grained, often temporary extension of it, letting a service like Bedrock decrypt on your behalf for a scoped set of operations. Because access to the plaintext runs through the unwrap call, editing the key policy or retiring a grant severs access to every artefact under that key at once, without touching the artefacts themselves. That is the revocation switch: the ciphertext stays exactly where it is and simply becomes unreadable.

The fourth thing is that cross-account and cross-service access is also just key-policy-and-grant work. If a log-analytics pipeline in another account has to read invocation logs, or a second account consumes a model artefact, the external principal must appear in the key policy or hold a grant, and its own IAM must allow the KMS actions. Both sides have to agree. There is no separate cross-account encryption feature to reach for; it is the same two levers pointed at an external principal.

Put together, every persistent artefact in the app reduces to the same four questions. Which artefact is this? Whose key encrypts it, AWS or yours? Do you need the control, or is the default fine? And can you audit and revoke, which is only true when the key is yours.

What we’ll filter on

  1. Whose key is it: AWS owned (invisible, no policy), AWS managed (visible, AWS-controlled policy), or customer-managed (your policy, your grants)?
  2. Can you audit use: does every decrypt show up in CloudTrail against a key you can inspect?
  3. Can you revoke independently: can you cut access by editing a policy or retiring a grant, without deleting or rebuilding the artefact?
  4. Does the artefact hold data or IP worth that control: your corpus, your model weights, your logs of real prompts, your agent’s session state?
  5. Who else needs in: same-account service principals only, or a cross-account principal that must be named in the key policy?

The landscape

Walk the artefacts a Bedrock app leaves at rest, and for each there is a place to attach a customer-managed key.

The Knowledge Base has several encryptable surfaces. The first is the source data. Documents usually sit in an S3 bucket, which defaults to SSE-S3 and can be switched to your own key through S3 default encryption, independent of Bedrock. Do that and the knowledge base service role needs kms:Decrypt on the key, scoped with a kms:ViaService condition for S3.

The second is the vector store holding the embeddings and the index. If you let Bedrock stand the store up for you, it passes a key you nominate through to Amazon OpenSearch Serverless or to Amazon S3 Vectors. An OpenSearch Serverless collection is always encrypted at rest, under an AWS owned key unless its encryption policy names a customer-managed one, and the key cannot be changed once the collection exists.

The third is the transient data written while a data source is ingested. The kmsKeyArn field on the data source puts that intermediate state under your key too. A fully managed knowledge base collapses these last two into one, taking a single key at creation that covers ingestion and the stored index together, with a grant Bedrock retires when the knowledge base is deleted. Retrieval sessions take a key of their own, through kmsKeyArn on a RetrieveAndGenerate request.

Custom and imported model artefacts are the intellectual-property case. Fine-tune a model or import your own weights and the resulting artefact is stored by AWS, under an AWS owned key by default. Both paths take a key of yours instead: customModelKmsKeyId on a customisation job, importedModelKmsKeyId on an import job. The AWS owned default encrypts the weights just as strongly. Only your key gives you the audit trail and the ability to cut access to them.

Model invocation logs are the record of what was actually asked and answered. Logging is off until you enable it, and it delivers to an S3 bucket, a CloudWatch Logs group, or both, in the same account and Region as the logging configuration. Neither destination uses a KMS key by default: S3 falls back to SSE-S3, and CloudWatch Logs encrypts log groups with its own server-side AES-GCM encryption unless you associate a key. Both take a customer-managed key, and for S3 the key policy has to let bedrock.amazonaws.com call kms:GenerateDataKey. This is often the most sensitive artefact of all, because it captures real prompts and completions with whatever customer data they quoted.

Agent session state persists the conversation context an agent carries across turns. Bedrock encrypts an agent’s information, control-plane data and session data alike, under an AWS owned key, and customerEncryptionKeyArn on the agent swaps in yours. Agents created before 22 January 2025 sit on an older arrangement, defaulting to an AWS managed key and needing their own pair of policies, so check which side of that date an agent was made on before writing the policy.

Then there is the supporting infrastructure that is not Bedrock-specific but is part of the same app. The S3 buckets holding source documents, exports, or staging data each take a customer-managed key. A Lambda backing an agent tool that stores anything, or whose environment variables hold configuration you want protected, can have those environment variables encrypted with a customer-managed key rather than the default AWS managed key. None of this is unique to generative AI; it is the ordinary KMS surface of the services the app is built from, and it belongs in the same key inventory.

Across all of these, transit is barely a decision. AWS requires TLS 1.2 for calls to Bedrock and recommends 1.3, and there is no unencrypted mode to fall into. Regulated workloads can point at a FIPS endpoint, but the baseline is already there. The choices worth making are all about the keys on the data at rest.

Evaluation

Side by side

Artefact Default key type Customer-managed key available You audit use (CloudTrail) You can revoke independently Typically holds
Knowledge Base source (S3) SSE-S3 ✓ ✓ ✓ Your corpus documents
Vector store / index AWS owned ✓ ✓ ✓ Embeddings, index
Custom / imported model AWS owned ✓ ✓ ✓ Model weights (your IP)
Invocation logs (S3 / CloudWatch) SSE-S3 / CloudWatch SSE ✓ ✓ ✓ Real prompts and completions
Agent session state AWS owned ✓ ✓ ✓ Live conversation context
Action-group S3 / Lambda env SSE-S3 / AWS managed ✓ ✓ ✓ Staging data, config

The table reads one way. For every artefact the default is already encrypted, and for every artefact a customer-managed key is available. The three right-hand columns are the same three every time: audit, independent revocation, and the control that comes with owning the policy. What differs down the rows is what the artefact holds, and therefore how much you want those three things. One difference does not show in the table. Some attachments are creation-time only, an OpenSearch Serverless collection and a fully managed knowledge base among them, so the key has to be chosen before the store exists rather than added to it later.

Customer-managed KMS key key policy + grants you own the policy CloudTrail: every decrypt disable / schedule deletion Knowledge Base source (S3) corpus documents, default encryption Vector store / index embeddings, ingestion-job state Custom / imported model model weights, your IP Invocation logs (S3 / CloudWatch) real prompts and completions Agent session state live conversation context Action-group S3 / Lambda env staging data, configuration one policy edit revokes access to all of them Envelope encryption: KMS wraps a per-artefact data key under this key; the bulk data is encrypted by the data key, and the unwrap call is where policy is enforced.
One customer-managed key can sit under many artefacts. Because every read routes through an unwrap call, the key policy is a single point of both audit and revocation.

The solution

The pick is customer-managed keys on the artefacts that hold data or IP, and a deliberate acceptance of the AWS owned default where none of that control is needed. The invocation logs, the Knowledge Base source and vector store, the custom or imported model, and the agent session state are the strong candidates, because each holds real customer content or your own model weights, and for each you want the audit trail and the revocation switch. The action-group buckets and Lambda environments come along for the same reason wherever they touch the same data. A short-lived staging bucket that holds nothing sensitive is a defensible place to leave the default; leaving it default is then a decision on the record, not an oversight.

Owning the key means owning the key policy, and that is where the control lives. The policy names the principals allowed to use the key, and for a Bedrock-managed artefact it has to let the Bedrock service principal decrypt on your behalf. Bedrock does that through grants. Attach a key to a custom model and it creates one long-lived primary grant, retired when the model is deleted, plus short-lived secondary grants for each asynchronous job. Narrow the permissions with an encryption-context condition on the resource ARN, or a kms:ViaService condition naming the Bedrock endpoint, so the grant only works where you meant it to. Tighten too far and the symptom is not a security alert: the service loses the ability to read its own artefact, and ingestion or logging fails with a KMS access error. Test the policy by confirming the artefact is still usable after the key is attached.

Revocation is the capability people underuse until an incident. Because access to plaintext runs through the KMS unwrap call, disabling the key or removing a principal from its policy makes every artefact under that key unreadable, without deleting or moving the data. The timing is worth being precise about. The key itself stops working almost at once, subject to eventual consistency, and a key-policy edit takes a short while to propagate through KMS. Data already protected by a data key is unaffected until that data key has to be unwrapped again, so a running job can keep reading for a while after the switch. That is still the response to a compromised principal or a contractual off-boarding: disable the key, and the corpus, the logs and the model go dark while you investigate, then re-enable to restore access. Scheduling deletion is the destructive version. The waiting period runs from 7 to 30 days, defaults to 30, and can be cancelled up to the moment it expires, after which ciphertext under that key is unrecoverable.

Audit is what compliance actually asked for. Every cryptographic operation against a customer-managed key is a CloudTrail event: which principal, which key, which encryption context, when. That turns “who read the corpus” and “what decrypted the invocation logs” into queryable history rather than a matter of trust. An AWS managed key gives you that much, since it is visible and CloudTrail-logged. What it does not give you is the policy control and the independent revocation, and an AWS owned key gives you none of the three.

Cross-account, when it appears, is the same two levers aimed outward. A log-analytics pipeline or a model-artefact consumer in another account has to be named in the key policy or hold a grant, and its own IAM has to permit the KMS actions. One useful limit: invocation logging itself will not deliver across an account boundary, so the cross-account reader reads a bucket in your account rather than a destination in theirs. The plainer you keep the set of principals on each key, the easier both the audit and the eventual revocation stay.

Worked example

The team creates a single customer-managed key for the sensitive data plane and puts four things under it: the S3 bucket holding the Knowledge Base documents, the OpenSearch Serverless collection backing the vector index, the S3 destination for model invocation logs, and the imported model artefact, whose weights were fine-tuned on the contract corpus. Two of those are set by editing the bucket’s default encryption. The model artefact takes the key ARN on the import job, and Bedrock creates a grant so it can decrypt for that resource only. The collection is the one they have to get right first time: its encryption policy names the key at creation, and a collection’s key cannot be changed afterwards.

Now the control is real and testable. On the audit side, a week later CloudTrail shows every decrypt against the key: the ingestion job reading the source bucket, the retrieval calls unwrapping index data, the logging pipeline writing completions. Compliance gets the “who read what, when” report from the key’s own event history rather than from a promise.

On the revocation side, a contractor’s role that had been granted use of the key is off-boarded. Removing that principal from the key policy is the whole action, and nothing has to be re-encrypted or moved. Their access ends once the edit has propagated and the next unwrap call fails. To rehearse a breach, the team disables the key in a staging copy and watches retrieval and logging fail as each cached data key runs out and the next unwrap call fails, then re-enables and watches them recover. The gap between throwing the switch and the last in-flight read stopping is the number worth knowing before an incident, not during one.

The envelope mechanics stay invisible through all of this. Each artefact has its own wrapped data key; the bulk contents were never encrypted directly under the KMS key, so attaching, auditing, and revoking never touched the gigabytes. One key, four artefacts, three proven capabilities: audit, reversible revocation, and a policy the team, not AWS, controls.

What’s worth remembering

  1. Everything is already encrypted at rest by default; what matters is whose key it is. Only a customer-managed key gives you the policy, the audit trail, and independent revocation.
  2. The key policy, extended by grants, is the gate on encrypted data. Editing it revokes access to every artefact under that key without touching the artefacts, though a data key already in use keeps working until the next unwrap.
  3. Put customer-managed keys on the artefacts that hold data or IP: Knowledge Base source and vector store, custom or imported models, invocation logs, agent session state, and the supporting S3 buckets and Lambda environments. Some of those, an OpenSearch Serverless collection among them, only take a key at creation.
  4. Disable a key for a reversible stop during an incident. Scheduling deletion runs a waiting period of 7 to 30 days, 30 by default, and after it expires the ciphertext under that key is gone for good.
  5. A too-tight key policy shows up as the service failing to read its own artefact, not as a security alert. Least privilege must still admit the Bedrock service principals that need to decrypt.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.