Series
Distributed Systems
What breaks when the machines are not the same machine, worked one pattern at a time in Unmesh Joshi's dependency order, so nothing here needs a pattern that has not been covered yet. The write-ahead log first, then replication and leadership, then the clocks that let separate machines agree what happened before what, then partitioning and the patterns for finishing work that spans several of them. Each post takes one pattern: what the mechanism actually is, where the cost goes, the shape of problem it fits, when to reach for something else, and the Go to write it. Sister to Concurrency, which covers one machine, and to Integration, which covers the messages between them. Part of Under the Hood.
The Write-Ahead Log
A write-ahead log is a file that changes are appended to before they are made, and the order of operations is what makes it work: frame the change, append it, fsync, then apply. The framing costs 67 nanoseconds an entry on the machine below and the fsync costs 216 microseconds, three thousand times more, which is why every real implementation groups entries behind one fsync: batching 128 of them cut the per-entry cost by a factor of 77. It also hands you a file whose tail is torn by every crash, so recovery has to treat the last entry as suspect, and 576 single-bit flips over a 72-byte log were all rejected by a four-byte checksum. What it does not do is bound its own size.
Coming soon