The situation
A parcel carrier runs 38 VPCs across 21 accounts, all attached to one transit gateway in eu-west-1, addressed out of 10.64.0.0/12. Two data centres carry the on-premises side, summarised as 10.16.0.0/13. Hybrid connectivity is three Direct Connect circuits and a VPN, and every one of them advertises that same /13 with the same customer ASN, 64901:
- Circuit A, 10 Gbps dedicated, cross-connected at a Direct Connect location whose associated Region is eu-west-1. One transit virtual interface.
- Circuit B, 10 Gbps dedicated, in the second data centre, cross-connected at a location whose associated Region is eu-west-2. One transit virtual interface.
- Circuit C, a 1 Gbps hosted connection at a small disaster-recovery site, at a second location associated with eu-west-1. One transit virtual interface. It was provisioned as a backup, and nobody expected it to move production traffic.
- Two Site-to-Site VPN tunnels to the same transit gateway, as last resort.
All three transit virtual interfaces terminate on a single Direct Connect gateway, associated with the transit gateway with allowed prefixes 10.64.0.0/12. No BGP community tags are set anywhere. No prefix lists filter anything.
The traffic looks like this at the evening peak. Out of the data centres towards AWS, roughly 2 Gbps on A and 2 Gbps on B, which is what everyone expected. Back from AWS towards the data centres, 9 Gbps total, split by equal-cost multi-path across A and C. Circuit A absorbs its half. Circuit C is a 1 Gbps link being handed 4.5 Gbps, and it drops what it cannot carry. That shows up as retransmits on the scan-event feed, and as a support queue full of drivers whose handheld terminals time out. Circuit B, ten gigabits on the invoice, carries almost nothing in that direction.
What actually matters
The first thing to be clear about is that a hybrid path has two directions and they are configured by two different parties. What you advertise to AWS decides how AWS returns traffic to you. What AWS advertises to you decides how you send traffic to AWS. Those are separate BGP conversations, and the attributes on one have no bearing on the other. A team that watches a link saturate and reaches for a route map, without first asking which direction is full, has a fifty per cent chance of moving the wrong dial. The change then stays in the configuration, because it looks like it should have worked.
The second is that a default is still a setting. When no local preference community tags are applied to a private or transit virtual interface, the paths are not treated equally. An AWS Region applies the medium local preference value 7224:7200 to advertisements from Direct Connect locations whose associated Region matches it, and something lower to the rest. Nobody here chose that, and it explains the whole picture. A and C sit at the same preference and form an ECMP set. B is a rung below, and it is selected only if both of the others withdraw. The circuit installed for resilience across two buildings carries almost nothing, because of a property of a colocation facility.
The third is that equal-cost multi-path is exactly what it says. It hashes flows across paths of equal cost, with no weighting for capacity, so a 10 Gbps circuit and a 1 Gbps circuit carrying the same prefix with the same attributes take the same share. There is no weighting control, so the design has to be expressed as preference, and the evaluation order that resolves preference is fixed. Longest prefix match runs first and outranks everything: a /14 with a long AS_PATH and low local preference still beats a /13 with every attribute set favourably. Local preference community tags run next, ahead of AS_PATH. AS_PATH length runs after that, and MED after AS_PATH, which is why AWS does not recommend leaning on it. A prepend applied to a path that already lost on local preference alters nothing, and it stays in the running configuration for two years, cited as evidence that BGP is unpredictable.
The fourth is that the number of prefixes is a design constraint rather than housekeeping. A transit or private virtual interface accepts 100 IPv4 prefixes per BGP session by default, raisable through prefix controls to 1,000 each for IPv4 and IPv6. A session handed more prefixes than its allocation goes idle and reports BGP status DOWN, so the circuit stops carrying traffic altogether. Going the other way, a transit gateway advertises 200 prefixes combined across IPv4 and IPv6 to on-premises over a transit virtual interface. Splitting a summary into more specifics to steer traffic uses part of that allocation permanently, on every session.
What we’ll filter on
Scoring each available knob against what the design needs:
- Steers the AWS-to-on-premises direction. Does setting it change which circuit returning traffic arrives on?
- Steers the on-premises-to-AWS direction. Does setting it change which circuit outbound traffic leaves on?
- Available on a transit or private virtual interface. The hybrid path to the 38 VPCs.
- Available on a public virtual interface. The path to S3 and the other public endpoints, which the same circuits also carry.
- Expresses strict active/passive. Can it hold the 1 Gbps link at zero until both 10 Gbps circuits are gone, rather than merely making it less attractive?
The landscape
Longest prefix match. Advertise 10.16.0.0/14 out of circuit A and 10.20.0.0/14 out of circuit B, keeping the covering 10.16.0.0/13 on both, and returning traffic follows the more specific route to whichever circuit carries it. It outranks every BGP attribute, so nothing set further down the chain can overrule it, and the covering /13 gives automatic failover: lose A, its /14 withdraws, and the /13 on B picks the traffic up. It takes three advertisements where there was one, on every session.
Local preference BGP community tags. Three values, 7224:7100 low, 7224:7200 medium, 7224:7300 high, applied by you to the prefixes you advertise, and honoured by AWS on private and transit virtual interfaces only. They are mutually exclusive, evaluated from lowest to highest with highest preferred, and evaluated ahead of AS_PATH. Tagging several connections with the same value is the documented way to load balance active/active, whether they are homed to the same Region or to different ones. If one of them fails, the rest load balance across what is left, regardless of home-Region association. Tagging one high and another low is the documented way to build active/passive.
AS_PATH prepending. Repeat your own ASN on the advertisement to make a path look longer and therefore worse. It works on both virtual interface types and it is the knob most people reach for first. On a private or transit virtual interface it is evaluated after local preference, so it only decides anything among paths that already tied there. On a public virtual interface, prepending is the mechanism, along with prefix length, since local preference communities do not apply. There is a catch on public virtual interfaces. If you peer with a private ASN, Direct Connect replaces it with 7224 when advertising your prefixes onward, and the prepends are stripped with it. Prepending steers nothing there unless you own and announce a public ASN.
MED. Read only when prefix length, local preference and AS_PATH have all tied. AWS does not recommend using it, given how far down the order it sits. One default is worth knowing. The transit gateway assigns a MED of 0 to inbound routes on Direct Connect attachments and 100 to inbound routes on VPN and Connect attachments, which is part of why the VPN stays in reserve with nobody configuring it.
Public virtual interface scope communities. 7224:9100 for the local Region, 7224:9200 for all Regions on a continent, 7224:9300 for all public Regions, applied to the prefixes you advertise, with global as the default if you tag nothing. They control how far into the Amazon network your prefix propagates, so they scope reachability rather than share load. Going the other way, Direct Connect tags what it advertises to you: 7224:8100 for routes originating in the Region associated with that point of presence, 7224:8200 for the same continent, and no tag for anywhere else. That is what you filter on if you want a public virtual interface to carry only local Regional traffic.
Your own routers. On the outbound direction, nothing AWS sends distinguishes one circuit from another. For a transit gateway association the Direct Connect gateway advertises exactly the allowed prefixes, originating from the Direct Connect gateway ASN, identically on every transit virtual interface hanging off it. There is one allowed-prefix list per association, so there is no AWS-side lever that makes circuit A’s advertisement look better than circuit B’s. Outbound path selection is decided entirely by local preference, weight and IGP metrics on your own equipment.
A note on what does not apply here. ECMP across Direct Connect gateway attachments forms only when the network prefix, prefix length and AS_PATH are exactly the same. The transit gateway does not support BGP Multipath AS-Path Relax, so paths carrying different ASNs do not form an ECMP set. A single Direct Connect gateway supports ECMP across multiple transit virtual interfaces, which is why all three circuits belong on one gateway rather than three.
Evaluation
Side by side
| Knob | Steers AWS to on-prem | Steers on-prem to AWS | Transit / private VIF | Public VIF | Expresses strict active/passive |
|---|---|---|---|---|---|
| Longest prefix match (split advertisement) | ✓ | ✗ | ✓ | ✓ | ✓ |
| Local preference communities (7224:71/72/7300) | ✓ | ✗ | ✓ | ✗ | ✓ |
| AS_PATH prepending | ✓ | ✗ | ✓ | ✓ | ✗ |
| MED | ✓ | ✗ | ✓ | ✗ | ✗ |
| Public VIF scope communities (7224:91/92/9300) | ✗ | ✗ | ✗ | ✓ | ✗ |
| Local preference / weight on your own routers | ✗ | ✓ | ✓ | ✓ | ✓ |
The second column has one tick in it, and that tick is the diagnosis. Five of the six knobs move returning traffic and only one moves outbound traffic. A team looking at a saturated inbound link therefore ends up editing the advertisement that governs it and calling the result a mystery. Prepending gets a cross in the last column because it is a nudge. Three prepends make a path worse, but a congested path is still a valid path, so BGP keeps using it, and a fourth circuit that also prepends can accidentally tie.
Which knob for which intent
The solution
Tag the two 10 Gbps circuits identically. Apply 7224:7200 to the 10.16.0.0/13 advertised out of circuit A and to the same prefix advertised out of circuit B. That replaces the implicit preference from the locations’ associated Regions with an explicit one that rates the two equally, and return traffic forms an ECMP set across both. Keep the AS_PATH identical on the two, keep ASN 64901 on both, and keep both transit virtual interfaces on the one Direct Connect gateway. ECMP forms only where prefix, prefix length and AS_PATH match, and AS-Path Relax is not available to bridge different ASNs. If A fails, B carries everything with no configuration change, which is the documented behaviour for connections tagged with the same community.
Tag the disaster-recovery circuit low. 7224:7100 on the /13 advertised out of circuit C puts it a rung below the pair, so it carries nothing while either 10 Gbps circuit is up and takes over when both are gone. This is the part the prepend could not do. Local preference is evaluated before AS_PATH, so the low tag settles the selection, while a prepend only separates paths that have already tied on local preference.
Leave the VPN alone. Two mechanisms already hold it in reserve. At the transit gateway, routes with the same CIDR from different attachment types are ranked by attachment, and Direct Connect gateway propagated routes sit above Site-to-Site VPN propagated routes. Separately, inbound routes on a Direct Connect attachment get a default MED of 0 while VPN and Connect attachments get 100. The transit gateway shows only the preferred route, so the VPN’s copy of 10.16.0.0/13 will not appear in the route table until the Direct Connect path stops being advertised. Worth knowing before somebody opens a support case about a missing route.
Split the advertisement where you want determinism rather than a hash. ECMP spreads flows, not bytes, and a hash can settle unevenly on a small number of large flows. If the requirement is that the northern data centre’s traffic returns on circuit B rather than roughly half of it, advertise 10.20.0.0/14 out of B and 10.16.0.0/14 out of A, and keep the covering /13 on both so a circuit loss reconverges without a human. Longest prefix match runs ahead of every attribute, so this is the strongest instrument available and the one that cannot be accidentally overridden. It also turns one prefix into three on each of three sessions, against an allocation that drops the session when it is exceeded, so the summarisation policy has to be a deliberate design with headroom rather than whatever the redistribution happens to emit this quarter.
Handle the public virtual interfaces separately, because none of the above applies to them. Local preference communities are honoured on private and transit virtual interfaces only. On the public side you have prefix length and AS_PATH, plus one trap: with a private ASN, Direct Connect substitutes 7224 for your ASN when advertising onward and takes your prepends with it. The scope communities are a different tool for a different job, controlling how far your prefix propagates rather than which circuit is selected. In the other direction, filter on 7224:8100 if you want a public virtual interface to accept only routes originating in the associated Region, which is how you stop a London public virtual interface pulling traffic destined for Sydney. One more constraint catches people. Direct Connect advertises public prefixes with a minimum AS_PATH length of 3, so if the same Amazon prefixes also reach your edge from a transit provider with a shorter path, your routers may select the internet instead of the circuit, and the fix is local policy rather than anything sent by AWS.
Fix the outbound direction on your own kit. It is already balanced here, so nothing needs doing, but it is worth understanding why nothing on the AWS side would have helped if it were not. The Direct Connect gateway advertises the allowed prefixes list, 10.64.0.0/12, from the Direct Connect gateway ASN, on every transit virtual interface associated with it. There is no per-circuit allowed-prefix list, and for a transit gateway association what you receive is precisely the allowed prefixes rather than the VPC CIDRs behind them. If you want outbound traffic pinned per data centre, set local preference or weight on your own routers, or tune the IGP metric between the two sites so each one exits locally. The same design that reaches the transit gateway across two Regions, described in the second-Region design, inherits this property: one allowed-prefix list per association, advertised identically everywhere.
Turn on BFD while you are in there. The default BGP hold timer on Direct Connect is 90 seconds with a 30 second keepalive, which is how long a failure that does not drop the physical link can blackhole traffic before BGP notices. Asynchronous BFD at the AWS liveness detection minimum interval of 300 ms, with a multiplier of 3, brings detection down to under a second. Do not configure graceful restart and BFD together; AWS recommends against it, and the 120 second graceful restart timer will hold a dead path open through exactly the failure BFD was added to catch.
Worked example
Take the single prefix 10.16.0.0/13 as it reaches the transit gateway, as found and under three changes.
As found
Circuit A carries it with local preference medium, applied implicitly because that location’s associated Region is eu-west-1. Circuit C carries it with local preference medium for the same reason. Circuit B carries it with something lower, because its location is associated with eu-west-2. The VPN carries it too, with a default MED of 100 and a lower attachment-type priority. Longest prefix match ties, all four being /13. Local preference resolves the three Direct Connect paths: A and C tie at the top and form an ECMP set, and B does not appear. The VPN is ranked below Direct Connect by attachment type, so its copy never appears either. Four and a half gigabits of return traffic goes down a one-gigabit hosted connection, and 10 Gbps of circuit B does nothing.
The prepend that changes nothing
An engineer adds three prepends of 64901 to the advertisement leaving the second data centre, reasoning that making circuit B look worse will push outbound traffic onto circuit A. Two things happen, and neither is the intended one. Outbound path selection is decided by what AWS advertises and by local policy, and AWS advertises the identical 10.64.0.0/12 on all three transit virtual interfaces, so the outbound split does not move by a single flow. On the return direction, where the prepend does apply, it changes nothing either: circuit B had already lost on local preference, which is evaluated before AS_PATH, so the prepend never gets read. The configuration now contains a change that is doing no work in either direction, and the next person to read it will assume it is doing something.
The prepend that only half works
A later attempt sets no communities and instead prepends twice on circuit C, to stop the hosted connection taking half the return traffic. That part works. A and C tie on local preference, so AS_PATH decides between them, and C drops out of the ECMP set. All nine gigabits now land on circuit A alone, because B still sits a rung below on the implicit local preference and never entered the comparison. Someone then prepends twice on circuit B as well, to keep the two data centres symmetric, and that changes nothing: B had already lost before AS_PATH was read. One 10 Gbps circuit is carrying everything, and four prepends in the configuration describe a policy nobody wrote down.
The configuration that holds
7224:7200 on A, 7224:7200 on B, 7224:7100 on C, no prepends anywhere, one ASN, one Direct Connect gateway, BFD enabled on all three virtual interfaces. Return traffic splits across A and B, C sits at zero, the VPN stays out of the route table until Direct Connect withdraws. Where the split needs to be deterministic instead of hashed, one /14 on each of A and B with the /13 kept on both as the covering route, and a written note of how many prefixes that leaves against the session allocation.
What’s worth remembering
- What you advertise to AWS steers return traffic, and the outbound direction is steered on your own routers, because every transit virtual interface receives the same allowed prefixes.
- Direct Connect evaluates longest prefix match, then local preference community tags, then AS_PATH, then MED, so a prepend on a path that already lost on local preference changes nothing.
- With no community tags applied, an AWS Region already sends return traffic to Direct Connect locations whose associated Region matches it, at medium local preference, which is why two circuits in different facilities do not load balance by default.
7224:7200on both prefixes is active/active;7224:7300on one and7224:7100on the other is active/passive; the tags are mutually exclusive and work only on private and transit virtual interfaces.- Public virtual interfaces get prefix length and AS_PATH for path selection,
7224:9100/9200/9300for how far a prefix propagates, and prepends that are stripped if you peer with a private ASN. - A transit or private virtual interface takes 100 IPv4 prefixes per BGP session by default, raisable to 1,000 with prefix controls, and exceeding the allocation drops the session rather than trimming the table.