Under the Hood · The Service Manual

How CloudTrail Actually Works

· 53 min read

Someone deleted a production bucket at 2am and you need to know who. The answer is sitting in CloudTrail, has been since the moment it happened, and finding it takes five minutes if you know how the records are shaped and five hours if you don’t. This is the service you configure once, ignore for a year, and then need more than any other.

The flight recorder

CloudTrail launched at re:Invent in November 2013, and the pitch was almost apologetically simple: a log of every API call made in your AWS account, delivered as JSON files to an S3 bucket. No dashboards, no alerting, no analysis. Just the record. The timing was no accident. By 2013, enterprises and governments were moving regulated workloads into AWS, and their auditors were asking a question AWS could not yet answer: who did what, when, and from where? Every compliance framework worth the name (PCI DSS, SOC 2, HIPAA, FedRAMP) requires an audit trail of administrative activity, and before CloudTrail there wasn’t one. You could see the current state of your account but not the sequence of actions that produced it.

The mental model that makes CloudTrail click is this: everything in AWS is an API call, and every control-plane API call is an event. When you click “Launch instance” in the console, the console calls RunInstances on your behalf. When Terraform creates a security group, that’s CreateSecurityGroup. When an EC2 instance’s role fetches a secret, that’s GetSecretValue. The console, the CLI, the SDKs, CloudFormation, and every third-party tool all funnel through the same public APIs, and CloudTrail sits behind those APIs recording each call as a structured event. There is no side door. If a resource changed, an API call changed it, and the call left a record.

That makes CloudTrail the flight recorder of your account. Like a flight recorder, it’s boring right up until the moment it’s the only thing that matters. Security incident response in AWS is substantially about reading CloudTrail well: reconstructing what an attacker did from the sequence of calls they made, working out which credentials were used, tracing an assumed role back to the human who assumed it. The service has grown a lot of surface since 2013 (data events, Insights, network activity events, a SQL-queryable lake) but the core has never changed: API call in, JSON record out.

One boundary to fix in your head early: CloudTrail records the control plane by default; the data plane is opt-in. Creating a bucket is a management event and gets recorded for free. Reading an object out of that bucket is a data event, and is recorded only if you’ve explicitly turned data events on, at a price. Many incident investigations founder on exactly this line, so we’ll come back to it more than once.

From API call to record on disk

The pipeline starts inside the service you called. When a request arrives at, say, the EC2 API endpoint, the service authenticates it (that’s IAM’s job), executes it, and emits an event describing the call to CloudTrail’s ingestion layer. The event is generated by the service that handled the request, which explains some of CloudTrail’s quirks: coverage and latency vary by service, because each service team owns its own event emission. Nearly every AWS service is integrated for management events; a growing catalogue supports data events; a handful of stragglers and preview features log nothing.

From ingestion, the event fans out to up to four destinations depending on what you’ve configured:

Event history gets every management event automatically, free, with no setup, retained for 90 days. It’s queryable in the console and via aws cloudtrail lookup-events, filtered by event name, resource name, username, and a few other attributes. It is per-region and management-events-only, but it is the reason a brand-new AWS account still has an audit trail: even if nobody ever configured anything, the last 90 days of control-plane activity are sitting there waiting.

Trails are the classic delivery mechanism: CloudTrail batches events and delivers them as compressed JSON files to an S3 bucket you own, under a predictable prefix (AWSLogs/{account-id}/CloudTrail/{region}/{year}/{month}/{day}/). Files land roughly every five minutes, each containing a batch of events, gzipped, typically a few kilobytes to a few megabytes. New trails are multi-region by default (this flipped in 2022; older accounts may still have single-region trails, which is almost never what you want). A trail can optionally fan events out to a CloudWatch Logs group as well, which gives you CloudWatch metric filters and alarms over the stream at a cost we’ll get to in the pricing section, because it surprises people.

CloudTrail Lake is the managed query layer: events flow into an event data store instead of (or as well as) S3 files, and you query them with SQL, with retention configurable up to ten years. More on Lake below, including the awkward fact of its status since mid-2026.

EventBridge receives management events (excluding read-only calls like Describe*, List*, and Get*) on the default event bus as they happen, wrapped in an AWS API Call via CloudTrail envelope. This path exists whether or not you have a trail, and it matters because it’s much faster than file delivery, which brings us to latency.

The honest answer on delivery latency is: minutes, not seconds, and occasionally worse. AWS documents that CloudTrail typically delivers events within about five minutes of the API call, and independent measurement backs that up: across a large sample, the average delay is around two and a half minutes in busy accounts, the 99th percentile just over five minutes, with a long tail of stragglers that can run to hours (fewer than one event in three thousand takes more than ten minutes, but the record delays run past sixteen hours). Latency also varies by service; EC2 events take roughly twice as long as S3 events to arrive. In quiet regions, batching means events tend to show up at almost exactly the five-minute mark. There is no SLA on any of this.

The practical consequence: if you’re building detection or automation that needs to react inside a minute (revoking leaked credentials, quarantining an instance), file delivery is the wrong trigger. Use the EventBridge path, which typically delivers management events in seconds to tens of seconds, and treat the S3 files as the durable record rather than the alerting signal. And when you’re doing incident response, remember the tail: “the event isn’t in the logs yet” and “the event never happened” look identical for the first few minutes, and in rare cases much longer.

Anatomy of an event

Every CloudTrail event is a JSON object with the same top-level shape, and fluency in that shape is most of the skill. Here’s a trimmed real-world example, a role deleting an S3 bucket:

{
  "eventVersion": "1.10",
  "eventTime": "2027-09-14T18:03:41Z",
  "eventSource": "s3.amazonaws.com",
  "eventName": "DeleteBucket",
  "awsRegion": "ap-southeast-2",
  "sourceIPAddress": "203.0.113.42",
  "userAgent": "aws-cli/2.17.0 md/awscrt#0.20.11",
  "userIdentity": {
    "type": "AssumedRole",
    "principalId": "AROAEXAMPLE123456789:jenkins-deploy-1892",
    "arn": "arn:aws:sts::111122223333:assumed-role/DeployRole/jenkins-deploy-1892",
    "accountId": "111122223333",
    "accessKeyId": "ASIAEXAMPLEKEY",
    "sessionContext": {
      "sessionIssuer": {
        "type": "Role",
        "arn": "arn:aws:iam::111122223333:role/DeployRole",
        "userName": "DeployRole"
      },
      "attributes": {
        "creationDate": "2027-09-14T17:58:02Z",
        "mfaAuthenticated": "false"
      }
    }
  },
  "requestParameters": { "bucketName": "orders-archive-prod" },
  "responseElements": null,
  "readOnly": false,
  "eventCategory": "Management",
  "managementEvent": true,
  "recipientAccountId": "111122223333",
  "eventID": "b62d40f5-6a3f-4b2e-9d7a-example",
  "requestID": "5C2AEXAMPLE"
}

eventSource and eventName tell you what happened: the service (as its API endpoint hostname) and the API action. Mostly these map one-to-one onto the public API, but not always. Console operations frequently generate flurries of Describe* calls you didn’t consciously make, some event names differ from the SDK method that triggered them, and a few services emit synthetic events (eventType: AwsServiceEvent) for things AWS did on your behalf, like a KMS key rotation. When hunting, search on the API name you’d call from the CLI first, then broaden.

userIdentity is the field that repays study, because it answers “who”, and “who” in AWS is layered. The type field sets the frame: Root (the account’s root user; any unexpected appearance of this is an incident), IAMUser (a long-lived user; accessKeyId starting AKIA means a long-lived key), AssumedRole (the overwhelmingly common case in well-run accounts), AWSService (an AWS service acting for you), AWSAccount (a caller from another account), IdentityCenterUser, FederatedUser, and occasionally Unknown (some console sign-in and billing events).

For AssumedRole, the identity is a session, and the record shows both halves. The arn names the assumed-role session, and its final segment is the role session name, chosen by whoever called AssumeRole. This is forensic gold when it’s set well: AWS IAM Identity Center sets it to the signed-in username, EC2 sets it to the instance ID, and good CI systems set it to a build identifier. sessionContext.sessionIssuer names the underlying role, and sessionContext.attributes tells you when the session was created and whether MFA was involved. The accessKeyId beginning ASIA marks temporary credentials; that exact string reappears in every call the session makes, which is how you group one session’s activity together and how you trace it back to the AssumeRole call that minted it. There’s also an optional sourceIdentity attribute, set at assume time and immutable through role chaining, which exists precisely so that the original human’s identity survives a chain of role assumptions; if you administer a multi-account setup and aren’t requiring it, start.

sourceIPAddress looks self-explanatory and is full of traps. For a direct call from a laptop or server over the public internet, it’s the caller’s public IP, as you’d hope. But when an AWS service makes the call on your behalf, the field contains the service’s DNS name instead: cloudformation.amazonaws.com when CloudFormation created the resource, config.amazonaws.com when Config recorded it. Some console-initiated actions show AWS Internal. Calls made through a VPC endpoint show the caller’s private IP (with a vpcEndpointId field alongside), which is more useful than the NAT gateway’s address you’d otherwise see, but useless if you don’t also know which VPC that private range belongs to. Never build an allow-list detection on sourceIPAddress without handling the service-name and private-IP cases; plenty of real detections have silently matched nothing for months because of this.

requestParameters and responseElements carry the call’s inputs and outputs. Request parameters are where the specifics live: which bucket, which instance type, which policy document. Response elements are null for most read-only calls and for many mutations; where present, they can contain the pieces you need to pivot, the standout example being AssumeRole, whose response includes the temporary accessKeyId it issued. Sensitive values are redacted (HIDDEN_DUE_TO_SECURITY_REASONS for console sign-in passwords, and secret material never appears), and very large request or response bodies get truncated past the event size limit. And two failure fields matter as much as the successes: errorCode and errorMessage. A burst of AccessDenied errors is the classic signature of an attacker enumerating what stolen credentials can do; an investigation that filters to successful calls only misses the reconnaissance.

The chain of digests

An audit log you can’t prove is intact is a story, not evidence. CloudTrail’s answer is log file integrity validation, which you enable per trail and should enable on every trail.

With validation on, CloudTrail delivers an hourly digest file to a separate prefix (AWSLogs/{account-id}/CloudTrail-Digest/...) alongside the log files. Each digest lists every log file delivered in the previous hour with its SHA-256 hash, and the digest itself is signed with SHA-256 with RSA using a private key held by CloudTrail (public keys are retrievable via the API, rotated regularly). Each digest also contains the signature of the previous digest, so the hour-by-hour digests form a hash chain: tamper with any log file and its hash won’t match the digest; tamper with a digest and its signature fails; delete a digest and the chain has a visible gap.

Validation is one command:

aws cloudtrail validate-logs \
  --trail-arn arn:aws:cloudtrail:ap-southeast-2:111122223333:trail/org-trail \
  --start-time 2027-09-14T00:00:00Z

It walks the chain and reports every file as valid, modified, or missing. Two honest caveats. First, the chain proves the files haven’t been altered since CloudTrail delivered them; it can’t conjure back events from a window where logging was disabled, so the chain complements, rather than replaces, controls that stop logging being turned off. Second, almost nobody runs validate-logs until the day it matters, at which point it matters enormously: it’s the difference between handing an auditor or a court a bucket of JSON and handing them a cryptographically verifiable record. Turn it on, and put the validation command in your incident-response runbook so the first responder establishes the record’s integrity before building a timeline on top of it.

The full surface today

The 2013 service was one thing; the 2027 service is six things sharing a console. Here’s the map, and when each is the right tool.

Event history (free, automatic, 90 days, management events only) is the right tool for “what just happened”: quick lookups during an incident’s first hour, checking what a colleague changed yesterday, small accounts with no formal logging setup. Its limits are the retention, the management-only scope, and clumsy filtering (one attribute at a time).

Trails are the durable backbone. The first copy of management events is free; you pay S3 storage (cheap, compressed, and lifecycle-manageable) and only pay CloudTrail itself if you add data events or extra copies. A trail is the right tool for long-term retention, for feeding a SIEM, and for Athena-based investigation. It’s also the compliance baseline: effectively every AWS security standard starts with “a multi-region trail exists, is logging, and delivers to a protected bucket”.

Data events extend the record into the data plane, per resource type, opt-in on a trail or event data store. The catalogue started with S3 object-level calls and Lambda Invoke and now covers dozens of resource types: DynamoDB item-level actions, S3 access points and directory buckets, EBS direct APIs, SNS publishes, SQS messages, Cognito, Bedrock model invocations, and a steadily growing list. Advanced event selectors are the control surface: field-level filters (on eventName, resources.ARN, readOnly, and so on) that let you log, say, only write operations on two sensitive buckets rather than every GetObject in the account. At data-plane volumes, writing selectors carefully is the difference between a useful signal and a bill-shaped firehose. Since late 2025 there’s also data event aggregation, which rolls data events up into five-minute summaries (top actions, error rates, access frequency) so you can watch access patterns without paying to store every individual event.

Network activity events, generally available since February 2025, are the newest category and close a real gap: API activity crossing your VPC endpoints, logged from the network’s point of view. Their headline value is deny visibility for data-perimeter work: when a VPC endpoint policy blocks a call (VpceAccessDenied), that denial previously vanished; now it’s an event. This is how you see someone inside your VPC trying to reach an S3 bucket in an account you don’t trust, or credentials from outside your organisation being used inside your network. Coverage is five services so far (S3, EC2, KMS, Secrets Manager, and CloudTrail itself), priced like data events, and since mid-2026 filterable by user identity in advanced selectors, so you can log only the denials from principals you don’t recognise. One sharp edge: for a denied caller from outside your account, the event deliberately carries minimal identity (no ARN or user agent, just a source IP, an account ID, and an opaque principal ID), because you’re not entitled to another account’s identity details.

Insights events are CloudTrail’s anomaly detection: enable them and the service builds a rolling statistical baseline of your management-event write activity, then emits an Insights event when call rates or error rates deviate from it. The two analysis types (ApiCallRateInsight and ApiErrorRateInsight) are enabled and charged separately, and since November 2025 Insights can also watch data-event patterns. Expect up to 36 hours after enabling before anything can fire, and set expectations accordingly: Insights is good at “this account is suddenly making 40x the usual number of RunInstances calls” (cryptomining, runaway automation, enumeration) and is not a threat-detection product. GuardDuty consumes CloudTrail with actual threat intelligence behind it; Insights tells you your account’s rhythm changed.

CloudTrail Lake deserves its own paragraphs, both for what it is and for what happened to it. Lake, launched in 2022, is a managed audit data lake: events flow into immutable event data stores with retention up to ten years, you query them with real SQL (joins, aggregations, nested fields), you get prebuilt and AI-assisted dashboards, you can federate an event data store into Athena, and you can ingest non-AWS events (your own applications, other clouds, SaaS audit logs) through channels, putting all your audit activity behind one query language. For organisations that never built a SIEM pipeline, Lake was the shortcut: ten years of queryable audit history with zero infrastructure.

The awkward fact: AWS closed CloudTrail Lake to new customers on 31 May 2026. Existing customers keep working (organisation-level event data stores continue to cover new accounts and regions; account-level stores continue for their existing accounts but won’t pick up new ones), but the feature now receives only critical fixes, and AWS’s guidance points new workloads at CloudWatch’s unified log management instead, which has grown OCSF normalisation, OpenSearch-powered analytics, and Apache Iceberg access to fill the role, plus a migration path that imports Lake event data stores into CloudWatch. Trails, event history, Insights, and the rest of CloudTrail are unaffected and fully supported. If you’re on Lake today, nothing forces you off it; if you’re designing fresh, the durable pattern is the one that predates Lake and will outlast it: a trail to S3, queried with Athena, or shipped to whatever SIEM you already operate. The event record is the stable thing; the query layers over it keep changing.

The edges

The edges of CloudTrail are where investigations go wrong, so they’re worth cataloguing.

Global services and the pull of us-east-1. IAM, STS, and CloudFront are global services whose control planes historically live in us-east-1, and their events are recorded there. A single-region trail in Sydney simply never sees CreateUser or AttachRolePolicy (and since 2021, single-region trails outside us-east-1 don’t receive global service events at all). Multi-region trails handle this correctly, which is one more reason they’re the default; the subtlety that survives is where to look: when you’re hunting IAM activity in event history or in the S3 prefix layout, it’s filed under us-east-1, not the region you work in. STS is the halfway case: calls to the global sts.amazonaws.com endpoint log in us-east-1, while calls to regional STS endpoints (which AWS has been nudging everyone towards for years) log in their own region. An AssumeRole event can therefore be sitting in a different region from the activity of the session it created. Budget your searches accordingly.

What CloudTrail does not capture. The list is longer than intuition suggests. Data-plane actions without data events enabled: if S3 data events weren’t on, nobody read, wrote, or deleted an object as far as the record is concerned. Anything below the API surface: an SSH or RDP session inside an EC2 instance is invisible (the StartSession call for Session Manager is logged, but not what happened inside; keystrokes need Session Manager’s own logging, and OS activity needs an agent on the host). Queries inside your RDS database. Traffic that isn’t an AWS API call at all, which is VPC Flow Logs’ territory. And CloudTrail records actions, not state: it tells you a security group was modified and by whom, but reconstructing what the rules looked like on 3 March is AWS Config’s job, which snapshots resource configuration over time and can replay its history. The trio is complementary, not overlapping: CloudTrail is who did what, Config is what things looked like, CloudWatch is how things behaved. Investigations usually need at least two of the three.

Redaction and truncation. Events have a size limit (256 KB classically; Lake accepts up to 1 MB since 2025), and oversized request or response bodies are dropped with a marker. Sensitive parameters are redacted. This is nearly always fine and occasionally maddening, most famously with PutRolePolicy-style calls where you want the full policy document that was attached and get it, versus large batch calls where the interesting item is in the part that got truncated.

Eventual and unordered delivery. Events within a log file aren’t sorted, events for one logical action can land in different files minutes apart, and eventTime has one-second granularity, so ordering two calls in the same second is guesswork. Build timelines on eventTime plus requestID correlation, not on file arrival order.

Duplicate-ish records across accounts. For cross-account activity, the event appears in both accounts’ trails with the same eventID but different recipientAccountId, and some fields visible to the resource owner differ from those visible to the caller, the network-activity deny case above being the extreme version. When two accounts’ records of one call disagree in detail, that’s the design.

The cost shape

CloudTrail’s pricing looks simple and hides two traps.

The simple part: your first copy of management events is free in every region, forever. S3 storage for the trail is your only cost, and compressed JSON audit logs are small; a mid-sized account’s management events cost pennies a month to store, and lifecycle rules to Glacier tiers make ten-year retention nearly free. There is no reason any AWS account should lack a trail.

The metered parts: additional copies of management events (a second trail logging the same events) cost USD$2.00 per 100,000 events. Data events cost USD$0.10 per 100,000, network activity events the same, Insights USD$0.35 per 100,000 events analysed (per analysis type, so both types doubles it), and data-event aggregation USD$0.03 per 100,000 analysed. Lake, for those grandfathered onto it, charges by ingested volume: USD$0.75/GB with one year’s retention included and about USD$0.023/GB-month to extend (up to ten years), or a tiered USD$2.50/GB option with seven years included.

Trap one is data-event volume. USD$0.10 per 100,000 sounds like nothing until you do the arithmetic on a busy data plane: a bucket handling 5,000 GETs a second generates around 13 billion events a month, which is roughly USD$13,000 in CloudTrail charges, plus the S3 storage for all that JSON, for a feature someone enabled with one checkbox “for visibility”. Data events on everything is almost never the right call. Advanced event selectors exist to make the checkbox precise: writes but not reads, these buckets but not those, deny errors but not successes. Decide what question you’d ask in an incident, log what answers it, and let aggregation cover the trend-watching.

Trap two is the CloudWatch Logs multiplier. Fanning a trail out to CloudWatch Logs re-prices your events as log ingestion, at USD$0.50/GB in the standard class, and verbose JSON adds up: the same management events that were free as a first trail copy can cost hundreds or thousands a month as CloudWatch ingestion in a busy account, before the log-group storage. Sometimes that’s worth it (metric filters and alarms on the stream are genuinely useful), but decide deliberately, and consider whether an EventBridge rule on the specific events you care about gets you the alerting without re-ingesting the entire firehose.

Running it in anger

A single account with a default trail is fine. An organisation needs a design, and the design has been stable for years.

One organization trail, delivered to a dedicated log-archive account. Created from the management account (or better, a delegated administrator), an organization trail logs every member account, including ones created next year, into one bucket. That bucket lives in a log-archive account whose entire job is to hold audit data: nearly nobody can log into it, nothing else runs in it, and the security team’s tooling reads from it. Member accounts cannot touch the bucket; a compromised workload account can’t reach the record of its own compromise.

Treat the log bucket as the crown jewel, because attackers do. The first competent move after gaining admin in an account is to blind the flight recorder, and there’s a short menu: StopLogging, DeleteTrail, UpdateTrail (pointing delivery somewhere useless), PutEventSelectors (silently excluding the events that matter), or deleting the delivered objects. Each defence maps to a menu item. Service control policies deny the trail-tampering APIs organisation-wide except for a break-glass role. The bucket policy admits cloudtrail.amazonaws.com writes (condition-scoped to your trail’s ARN, so another account’s trail can’t write into your bucket) and almost nothing else. S3 Object Lock in compliance mode on the bucket makes delivered logs undeletable for the retention period, by anyone, including root in the bucket’s own account. Log file validation gives you the digest chain to prove what’s there is intact. And an EventBridge rule alerting on StopLogging, DeleteTrail, UpdateTrail, and PutEventSelectors means that even if tampering succeeds, the tampering call itself was logged and someone’s phone buzzed; GuardDuty raises a finding for trail-disabling out of the box.

Encrypt with SSE-KMS, carefully. Trail delivery supports SSE-KMS with a customer-managed key, which adds a second authorisation gate to reading logs: bucket access alone isn’t enough without kms:Decrypt on the key. That’s real defence in depth, and also a classic self-inflicted outage: scope the key policy wrong and you lock your own security tooling (or Athena) out of the logs. Keep the key in the log-archive account, next to the bucket.

Make the archive queryable before you need it. Raw trail files are thousands of small gzipped objects; the standard move is an Athena table over the bucket with partition projection on account, region, and date, so five years of organisation-wide logs are one SQL query away. Set it up before you need it, not during an incident. A query you’ll want ready:

SELECT eventtime, eventname, useridentity.arn,
       sourceipaddress, errorcode
FROM cloudtrail_logs
WHERE eventname LIKE 'Delete%'
  AND account = '111122223333'
  AND region = 'ap-southeast-2'
  AND date_parse(eventtime, '%Y-%m-%dT%H:%i:%sZ')
      > now() - interval '7' day
ORDER BY eventtime DESC;

Feed the SIEM from the bucket, not instead of it. If you run Splunk, Sentinel, or an OpenSearch pipeline, ship events from the log-archive bucket (S3 notifications into the ingest pipeline) and keep the bucket as the immutable source of record with longer retention than the SIEM’s hot storage. The SIEM answers this quarter’s questions; the bucket answers the auditor’s and the lawyer’s.

Who deleted that bucket

It’s Monday; orders-archive-prod is gone; nobody admits anything. Here’s the read.

Step one: find the event. Inside 90 days, event history answers directly:

aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=EventName,AttributeValue=DeleteBucket \
  --region ap-southeast-2

One hit: the event from the anatomy section above. eventTime 18:03 UTC Sunday, eventSource s3.amazonaws.com, and immediately one useful negative: errorCode is absent, so the call succeeded on the first try. No fumbling, which suggests the caller had done it before or scripted it.

Step two: read the identity, in layers. userIdentity.type is AssumedRole. The lazy read stops at sessionIssuer: “DeployRole did it”, which is true and useless, because forty humans and three systems can assume DeployRole. The full ARN ends /DeployRole/jenkins-deploy-1892: the role session name claims this was Jenkins build 1892. But session names are caller-chosen strings, not verified facts, so check the claim against the rest of the record. sourceIPAddress is 203.0.113.42; if that’s the CI runners’ egress IP, the story holds. If it’s residential broadband, someone else was driving the role and chose that session name to blend in. userAgent says aws-cli, where this pipeline normally speaks Terraform: a thread worth pulling. mfaAuthenticated false is expected for machine credentials.

Step three: trace the session to its origin. The event’s accessKeyId (ASIA...) names the temporary credentials. Search for the AssumeRole event whose responseElements.credentials.accessKeyId matches, remembering the regional wrinkle: STS events log in us-east-1 if the global endpoint was used, or in-region otherwise, so search both (this is where the Athena table beats clicking around event history). The matching AssumeRole event, at 17:58, shows its own caller in userIdentity: perhaps the EC2 instance profile of a Jenkins runner (instance ID as session name; story confirmed), or perhaps an IAM user’s long-lived AKIA key from an unfamiliar IP, at which point this stops being an ops retro and becomes credential-compromise response. If sourceIdentity was stamped at federation time, it names the human outright and the chain is one hop shorter.

Step four: mind the gap. DeleteBucket only succeeds on an empty bucket, so who emptied it? Those were DeleteObject calls, which are data events, recorded only if data events were enabled for that bucket. If they were, the same session’s ASIA key shows a burst of deletions in the minutes before 18:03 and the timeline is complete. If they weren’t, the emptying is simply invisible, and that silence, in a bucket that mattered, is the finding for the post-incident review: sensitive buckets get write data events, decided in advance, priced deliberately.

The wisdom that generalises: keep the free trail everywhere and never let anyone turn it off; enable data events on the resources you’d grieve, not on everything; make session names and sourceIdentity carry real provenance, because a session name is only as honest as the thing that set it; treat identity fields as claims to corroborate against IP, user agent, and time-of-day; remember the record is minutes behind reality and IAM’s part of it lives in us-east-1; and do the five minutes of Athena setup before the 2am call, because the flight recorder only pays off if you can read it under pressure.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.