Resilience

Nygard's stability patterns and the antipatterns they defend against, worked one at a time in Go. Every pattern here trades capacity, latency or correctness for survival, so each post says what it gives up as plainly as what it buys back. The primitives first, timeout and retry and fail fast, then the patterns that compose them, then the failure modes the whole set exists to prevent. Each post takes one entry: the mechanism, where the cost goes, the shape of problem it fits, when it is the wrong answer and what to do instead, and the Go to write it, with at least one diagram showing the thing going wrong rather than the thing working, because a pattern only makes sense against the failure it prevents. Sister to Distributed Systems and Concurrency. Part of Under the Hood.

What Resilience Actually Costs

Every one of the sixteen stability patterns gives something up: capacity, latency, or whether the answer is current. In a tick-level simulation of one caller tier in front of one dependency, 54,097 requests through an unguarded system lost 4,476 that needed nothing from the dependency at all; bounding the dependency to twenty concurrent calls took that number to zero and refused 12,053 outright. The same bound, with nothing wrong anywhere, refused 396 requests the unguarded system answered. Steady state, fail fast, and the difference between a system that degrades and one that stops are the vocabulary the rest of this run uses. Every later post answers the question stated once here: what does this one give up.

Coming soon