Exam Room · Advanced Networking Specialist

Keeping S3 Traffic Off the Internet in Forty Accounts

· 38 min read

Advanced Networking · part of The Exam Room

The situation

Marlowe Mutual writes household and motor policies. Its AWS estate is forty-one accounts in one Organization: thirty-eight of them hold a workload VPC in eu-west-1, one is the network account that owns the transit gateway, and two are security and logging. Every workload VPC has private subnets in three Availability Zones and a NAT gateway in each, which is a hundred and fourteen NAT gateways. The applications write claim documents to S3, read policy records from DynamoDB, fetch database credentials from Secrets Manager, and call KMS on nearly every request. All of that traffic leaves through the NAT gateways.

A data-loss-prevention review sampled a month of flow logs and CloudTrail. Around 215 TB a month goes to S3 address ranges, about 14 TB to DynamoDB, and roughly 90 million Secrets Manager and KMS calls carry almost no bytes at all. The written finding is not about volume. It is that every workload holds an unconstrained outbound path to any bucket in any account anywhere, including accounts nobody at Marlowe owns, and that nothing in the network stops a compromised task using it. The remediation date is the end of the quarter.

The complication is on the other side of the estate. Two data centres, one outside Reading and one outside Warrington, connect over a pair of 10 Gbps dedicated Direct Connect connections, transit virtual interfaces into a Direct Connect gateway associated with the same transit gateway. A nightly batch job on the policy administration platform writes about 900 GB of extracts into three of the same buckets, and a settlement service reads two secrets before every run. Whatever replaces the NAT path has to keep serving both.

What actually matters

Start with what the finding actually says, because the phrase “off the internet” is doing two jobs and only one of them is a routing problem. Traffic from a private subnet through a NAT gateway to S3 is already TLS-encrypted, and AWS is explicit that it stays on the AWS network even when it passes an internet gateway on the way. What the review objects to is that the path is unconstrained: the same route that reaches the company’s own buckets reaches every other bucket in the world. A remediation worth the quarter constrains two things at once. The route has to reach the service and nothing else, and the request that travels it has to be denied when it names a resource outside the organisation. Only the first is a network control. Fix the route, leave the authorisation open, and the finding closes on paper while the exfiltration path is still there, one bucket name away.

The second thing that decides this is who can use each mechanism. A route table entry serves the subnets associated with that route table, and nothing else in the world. Packets arriving on a transit gateway attachment, a peering connection, or a Direct Connect gateway are delivered to a destination address; the VPC’s route table is not applied to them, so a service-prefix route in it has no effect on that traffic. So the two data centres are not a detail to bolt on at the end. They rule out the unbilled answer for part of the traffic, and any mechanism serving a client from outside the VPC must present an address inside the VPC.

Then the money, which has two meters pointing in opposite directions. One is per endpoint, per Availability Zone, per hour, and it is multiplied by the number of services and the number of VPCs, so it grows with every repeat of the same endpoint across thirty-eight VPCs. The other is per gigabyte, and centralising makes it worse rather than better, because a shared hub adds a per-gigabyte transit gateway processing charge to traffic that would otherwise never have crossed it. The 215 TB of claim documents and the 90 million near-empty KMS calls sit on opposite ends of that trade. Sorting the estate by which meter each flow lands on is most of the design.

Last, ownership and the shape of a bad day. An endpoint in a workload VPC fails on its own and carries a policy the account team can write for its own workload. One shared set is a single thing to run, a single thing to lose, and a single policy document, capped at 20,480 characters, that has to be simultaneously correct for thirty-eight accounts. Sharing also makes DNS a dependency of the whole estate rather than of one VPC: once a service name resolves to an address in somebody else’s VPC, a resolution failure is an outage in thirty-eight accounts at once, and every newly vended account is broken until an association is made for it. That is a real operational cost set against a real saving, so it belongs in the comparison rather than in a footnote. The estate already runs on a hub, and a hub is where shared plumbing tends to accumulate.

What we’ll filter on

  1. Serves traffic that originates in the workload VPC’s own subnets.
  2. Serves traffic arriving from the data centres over Direct Connect, and from other VPCs over the transit gateway.
  3. Cost shape across thirty-eight VPCs and four services, counting both the per-endpoint-hour meter and the per-gigabyte meter.
  4. Where the authorisation control lives, and how many accounts share one policy document.
  5. Whether clients or resolvers have to be reconfigured, or whether the existing SDK call keeps working untouched.

The landscape

Four shapes are available for private access to AWS service APIs, and they are not variations on one idea.

A gateway endpoint exists for S3 and DynamoDB and for nothing else, and it is not PrivateLink. Creating one adds a route to the VPC route tables you select whose destination is an AWS-managed prefix list for the service and whose target is the endpoint. Longest prefix match means that route wins over the 0.0.0.0/0 entry pointing at the NAT gateway, so instances reach the service without one. There is no network interface, no address consumed from your subnets, no hourly charge and no per-gigabyte charge. Instances still address the service’s public IP ranges and still use the ordinary regional DNS name, which is why the security group needs an outbound rule referencing the prefix list, and why any network ACL needs the actual CIDR blocks from that prefix list, since ACLs cannot reference prefix lists. The account quota is 20 gateway endpoints per Region, adjustable, with up to 255 per VPC; at two per account nobody is close. What a gateway endpoint cannot do is stated flatly in both the S3 and the DynamoDB documentation: no access from on premises, and no access from another Region. Traffic arriving over peering or a transit gateway fails for the same reason: the VPC route table holding the endpoint route is not applied to it.

An interface endpoint is PrivateLink. It creates one requester-managed network interface per subnet you select, each with a private address from your own range, and the service is then reachable at an address inside the VPC. That is what makes it work from the data centres over Direct Connect, from peered VPCs, and from spokes over the transit gateway. S3, DynamoDB, Secrets Manager and KMS all support one. Billing, at US East list prices, is USD$0.01 per endpoint per Availability Zone per hour, plus data processing at USD$0.01 per GB for the first petabyte a Region a month, falling to USD$0.006 and then USD$0.004 above that. With private DNS enabled, the regional service name resolves to the endpoint’s addresses inside that VPC, with no application change. The default quota is 50 interface and Gateway Load Balancer endpoints per VPC, combined.

Centralised interface endpoints put one set in a shared-services VPC in the network account and let every spoke reach them over the transit gateway. Thirty-eight copies of an hourly charge collapse to one, and in exchange the traffic crosses the transit gateway, which adds USD$0.02 per GB of transit gateway processing on top of the endpoint’s own per-gigabyte charge. The catch is DNS. Enabling private DNS on an interface endpoint creates an AWS-managed private hosted zone that resolves only inside the VPC holding the endpoint, so it does nothing for a spoke.

Finally, keep the NAT gateways and filter the egress, allow-listing approved domains in front of them. It constrains where traffic may go without touching a single route. It also keeps a hundred and fourteen NAT gateways in the estate and keeps all 229 TB on the NAT meter, which at the US East (Ohio) rate of USD$0.045 an hour and USD$0.045 a gigabyte is roughly USD$3,700 a month in gateway hours and USD$10,300 a month in processing before anything else is counted. Ireland differs a little; the shape does not.

Evaluation

Side by side

Option In-VPC traffic From the data centres or the hub Cost at 38 VPCs Policy per account Nothing to reconfigure
Gateway endpoint in each VPC ✓ ✗ ✓ ✓ ✓
Interface endpoint in each VPC ✓ ✓ ✗ ✓ ✓
Centralised interface endpoints ✓ ✓ ✓ ✗ ✗
NAT gateways with an egress allow-list ✓ ✓ ✗ ✗ ✓

No row scores five, and the two rows that come closest fail on different filters. Gateway endpoints carry no charge, need no DNS work and leave the policy in each account’s hands, and they cannot serve a packet that started in Warrington. Per-VPC interface endpoints serve everything and bill for it: four services across three Availability Zones in thirty-eight VPCs is 456 network interfaces at USD$0.01 an hour, about USD$3,300 a month for a capability two of those four services already provide at no charge. Centralised endpoints fix the arithmetic and introduce a single shared policy document and a DNS dependency. Egress filtering never takes the traffic off the NAT path, which is what the finding was about.

Split the estate by flow instead of picking a row, and every filter is satisfied by something.

Where each flow lands

THE FLOW THE GATE WHERE IT LANDS In-VPC to S3, DynamoDB ~215 TB a month, 38 VPCs bulk claim documents In-VPC to KMS, Secrets Mgr ~90M calls a month almost no bytes Data centres to S3 ~900 GB nightly extracts over Direct Connect Data centres to DynamoDB settlement lookups over Direct Connect Starts in a route table this account owns? Service takes a gateway endpoint? Can a hosted zone override the service name? Gateway endpoint, all 38 VPCs prefix-list route, no ENI, no hourly or per-GB charge endpoint policy written per account security group must allow the prefix list outbound Centralised interface endpoints shared-services VPC, 3 AZs, reached over the hub private DNS OFF, one hosted zone per service name associated with all 38 consuming VPCs S3 interface endpoint, inbound only private DNS on, restricted to the inbound Resolver requires an S3 gateway endpoint in the same VPC on-prem forwarders point at the Resolver addresses DynamoDB endpoint-specific URL client configured with the vpce hostname directly no private hosted zone: AWS documents it as unsupported 50,000 requests per second per endpoint
Most of the answer follows from the first gate: traffic that starts in a route table the account owns can use a free gateway endpoint, and traffic arriving from the data centres cannot.

The solution

Gateway endpoints carry the in-VPC S3 and DynamoDB traffic, a centralised set of interface endpoints in a shared-services VPC carries everything else including the hybrid flows, and the two services with awkward DNS get handled individually.

Gateway endpoints for the bulk

Every one of the thirty-eight workload accounts gets an S3 gateway endpoint and a DynamoDB gateway endpoint, associated with each private subnet route table. That takes 229 TB a month off the NAT meter at no charge, since gateway endpoints have no hourly and no per-gigabyte charge. It also keeps that volume off the transit gateway. Centralising the same 215 TB would have added USD$0.02 a gigabyte of transit gateway processing and USD$0.01 a gigabyte of endpoint processing, about USD$6,500 a month on a path that is otherwise unbilled.

Two mechanical details bite during rollout. Security groups on the workload tasks need an outbound rule whose destination is the AWS-managed prefix list for the service, because instances still address the service’s public ranges. Any network ACL in the path needs the CIDR blocks written out, since ACLs cannot reference a prefix list, and those blocks change, so the ACL is the part that breaks two months later with nothing logged to say why. The NAT gateways stay for genuine internet egress such as package mirrors and third-party APIs; what collapses is the volume they process, not their existence.

One shared-services VPC for the rest

Secrets Manager and KMS have no gateway endpoint, so those 90 million small calls need PrivateLink. Doing that per VPC is 38 accounts times two services times three Availability Zones, 228 network interfaces, roughly USD$1,660 a month to carry a few hundred gigabytes. Centralising it is six network interfaces in the network account, about USD$44 a month, plus the transit gateway charges on traffic that is nearly all headers.

That is the right trade here, and it has consequences. The endpoint policy on a shared endpoint is one document, capped at 20,480 characters, that has to be right for every account using it, and the whole estate loses the service together if it breaks. Write those policies with organisation-scoped conditions rather than enumerating principals, or the character limit arrives faster than expected. Give each condition key its own statement: multiple keys inside one condition block are combined with AND, so a single statement naming both would deny only the request that fails both tests at once.

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyPrincipalsOutsideTheOrganisation",
      "Effect": "Deny",
      "Principal": "*",
      "Action": "*",
      "Resource": "*",
      "Condition": {
        "StringNotEquals": { "aws:PrincipalOrgID": "o-marlowe1234" }
      }
    },
    {
      "Sid": "DenyResourcesOutsideTheOrganisation",
      "Effect": "Deny",
      "Principal": "*",
      "Action": "*",
      "Resource": "*",
      "Condition": {
        "StringNotEquals": { "aws:ResourceOrgID": "o-marlowe1234" }
      }
    }
  ]
}

The same pair belongs on the per-account gateway endpoints, where maintenance is light because each account writes its own and can add a resource-level allow list on top of it. That is the half of the finding the routing change does not address: a gateway endpoint restricts the path to S3, and only the policy restricts which S3.

The DNS wrinkle that makes centralising work

Private DNS on an interface endpoint creates an AWS-managed private hosted zone visible only inside the VPC that holds the endpoint. A spoke VPC querying secretsmanager.eu-west-1.amazonaws.com gets the public answer and routes to its NAT gateway, which is where the whole exercise started. The fix is to turn private DNS off on the shared endpoint and build the resolution by hand: a Route 53 private hosted zone named for the service, an A record for the full service endpoint name aliased to the interface endpoint, and an association with every consuming VPC. Set Evaluate target health to No, which is what Route 53 documents for an alias whose target is a VPC interface endpoint.

The association is cross-account, so the console cannot do it. The hosted zone’s account authorises, then the VPC’s account associates, then the authorisation is deleted:

aws route53 create-vpc-association-authorization \
  --hosted-zone-id Z0123456789ABCDEFGHIJ \
  --vpc VPCRegion=eu-west-1,VPCId=vpc-0a1b2c3d4e5f67890

aws route53 associate-vpc-with-hosted-zone \
  --hosted-zone-id Z0123456789ABCDEFGHIJ \
  --vpc VPCRegion=eu-west-1,VPCId=vpc-0a1b2c3d4e5f67890

The quota to watch is 300 VPC associations per hosted zone. Thirty-eight is comfortable, and an estate heading past three hundred should move to Route 53 Profiles, which carry up to 5,000 private hosted zones and 1,000 VPC associations and turn this from a per-account association into a shared bundle. Put the association step into the account-vending pipeline on day one; the failure mode otherwise is a new account whose workloads fall back to the public endpoint with nothing to flag it, which is the exact condition the review found.

For the data centres, name resolution needs a Route 53 Resolver inbound endpoint in the shared-services VPC, two IP addresses at USD$0.125 per network interface an hour, with conditional forwarders on the on-premises resolvers for the service names. Each IP address in the endpoint takes up to 10,000 UDP queries a second. That figure falls as low as 1,500 when connection tracking is enforced by restrictive security group rules, or when queries arrive through a Network Load Balancer, so size on the lower number.

S3 and DynamoDB from the data centres

These two do not follow the pattern.

S3 has a purpose-built option. Create an S3 interface endpoint in the shared-services VPC, enable private DNS on it, and leave PrivateDnsOnlyForInboundResolverEndpoint set to true, which is the default; the call returns an error unless an S3 gateway endpoint already exists in that VPC. Regional S3 names then resolve to the interface endpoint only for queries arriving through the inbound Resolver endpoint, meaning the nightly batch job in Warrington reaches S3 privately with no client change, while in-VPC traffic in that same VPC keeps using the free gateway endpoint. Spoke VPCs are untouched: they resolve through their own Resolver, get the public answer, and hit their own gateway endpoints. Where a client can be reconfigured, the endpoint-specific name works with no hosted zone and no resolver at all, since bucket.vpce-0a1b2c3d4e5f-a1b2.s3.eu-west-1.vpce.amazonaws.com resolves from public DNS to the endpoint’s private addresses.

DynamoDB is the one to remember, because the obvious move is documented as wrong. AWS PrivateLink for DynamoDB does not support private or hybrid DNS, and the documentation carries an explicit instruction not to create private hosted zones overriding dynamodb.eu-west-1.amazonaws.com, because the service’s DNS configuration changes over time and a stale override sends requests back out over public addresses. Hybrid DynamoDB access means configuring the client with the endpoint-specific URL, https://vpce-0a1b2c3d4e5f-a1b2.dynamodb.eu-west-1.vpce.amazonaws.com, and each such endpoint takes up to 50,000 requests a second. At Marlowe the settlement service only reads secrets from on premises and never touches DynamoDB directly, so DynamoDB stays gateway-only and the problem does not arise. On an estate where it does, that client change is the work, and finding it in week eleven of a thirteen-week remediation is how these programmes slip.

Proving it, and the bucket side

Network controls answer where the packet went. The audit evidence the review wants is the other half: bucket policies denying any request whose aws:SourceVpce is not one of the approved endpoint IDs, or more maintainably a Deny on aws:ResourceOrgID mismatch at the endpoint. Bucket-side endpoint conditions lock out console access to the same bucket, since console requests do not arrive through the endpoint, so scope them to the paths the applications use and expect a ticket from whoever tries to download an object by hand. Flow logs in the workload VPCs stop showing NAT traffic to S3 ranges, which is the before-and-after the auditor can read without understanding any of this.

Worked example

A claims task writes a 4 MB document

The task in account 17 calls PutObject against marlowe-claims-eu-west-1. It resolves s3.eu-west-1.amazonaws.com through its own VPC Resolver and gets a public S3 address. Longest prefix match in the subnet’s route table picks the prefix-list route to the gateway endpoint over the 0.0.0.0/0 route to the NAT gateway. The endpoint policy denies anything outside the organisation, the bucket policy denies anything not arriving through one of the thirty-eight approved endpoint IDs, and the object lands. Nothing is billed for the path.

The same task fetches a database credential

GetSecretValue resolves secretsmanager.eu-west-1.amazonaws.com through the VPC Resolver, which consults the manually created private hosted zone associated with this VPC and returns a private address in the shared-services VPC. The packet crosses the transit gateway, hits the interface endpoint’s network interface, and PrivateLink carries it to Secrets Manager. Billed: transit gateway processing, endpoint processing, both on a payload of a few kilobytes.

The nightly extract from Warrington

The batch host resolves s3.eu-west-1.amazonaws.com against the on-premises resolver, which conditionally forwards to the Resolver inbound endpoint addresses. Because the S3 interface endpoint has private DNS restricted to the inbound Resolver, that query returns the endpoint’s private addresses rather than public ones. The upload crosses Direct Connect, the Direct Connect gateway and the transit gateway to the shared-services VPC, and PrivateLink delivers it to the same bucket the claims task wrote to an hour earlier. No change was made on the batch host.

What’s worth remembering

  1. Gateway endpoints for S3 and DynamoDB are route table entries with no hourly or per-gigabyte charge, and they serve only the VPC whose route table holds the entry, never on-premises, peering or transit gateway traffic.
  2. Interface endpoints put a private address inside the VPC, which is what makes them reachable from Direct Connect, peering and a transit gateway.
  3. Centralising interface endpoints means disabling private DNS and building a Route 53 private hosted zone per service name, associated with every consuming VPC, up to 300 associations per zone.
  4. PrivateLink for DynamoDB does not support private or hybrid DNS, so clients outside the VPC must use the endpoint-specific vpce hostname rather than a hosted zone override.
  5. An S3 interface endpoint can restrict private DNS to inbound Resolver queries, which serves on-premises clients while in-VPC traffic stays on the gateway endpoint that option requires in the same VPC.
  6. Changing the route satisfies half the finding; the endpoint policy and the bucket policy condition are what stop a workload reaching a bucket that is not yours.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.