System Design103 min total · 14 parts
System Design Fundamentals for Interviews: Scalability, Trade-offs, and the Framework Interviewers Actually Grade
Part 7 of 14 · ~5 min
Consistency Models and the CAP Theorem
Replicas, shards, caches — every scaling technique from the last two chapters worked by keeping extra copies of the same data lying around, and each of those copies opens up the identical question: for however briefly they disagree with each other, what does the system actually commit to promising you?
Strong vs. Eventual Consistency
Strong consistency is the promise that the instant a write finishes, every read that comes after — from any node, immediately — reflects it, full stop. No gap exists where a client could glimpse something stale; running a transaction against one single leader gets you precisely that guarantee.
Eventual consistency asks for less up front: it only guarantees that, assuming nothing else gets written in the meantime, every replica will eventually agree on the same value — with zero commitment on how long that catching-up actually takes. A read that happens to land on a replica still working through the backlog can hand back something out of date in the meantime. The lagging follower from the previous chapter, still mid-replication when a customer's refresh hits it, is precisely this in practice.
What you're trading is latency and availability against a window of staleness that's usually short but real. Strong consistency needs coordination with multiple nodes (or one authoritative one) before a write, or even a read, counts as finished — which takes time, and simply can't happen if those nodes aren't reachable. Eventual consistency lets whichever replica you can reach answer right now, at the cost of it occasionally lagging a beat behind.
Fanline learns exactly where that line needs to sit during the roughest ninety seconds of the Marlow Vance on-sale. Two different numbers on the site both say "seats remaining," and in the earliest build, both come from the same code path: a quick read off a nearby replica. For the browse page's little badge — "6 left at The Colosseum" — that's completely harmless; a number that's a few seconds behind costs nothing, and it refreshes on its own soon enough. But the reserve button, the one that actually puts a specific seat on hold for checkout, was reading availability off that same lagging replica before writing the hold. In the worst stretch of the on-sale, two fans in two different cities both saw the venue's last floor seat as open — because the replica each of them hit hadn't yet caught up to a reservation the leader had already accepted seconds earlier — and both walked away thinking they'd gotten it. One of them found out he hadn't at the door.
The fix wasn't making everything strongly consistent — that would mean every browse-page badge on the whole site paying a coordination cost it never needed. It was splitting the two paths on purpose: the reserve button reads and writes against the leader inside one transaction, so a single seat can never look available to two reservation attempts at once. The badge everyone else sees keeps pulling from a replica, because a browse-page number running a few seconds optimistic costs nothing, while a hold being wrong costs an awkward phone call and a comped upgrade.
What CAP Actually Claims
The CAP theorem says a distributed data store can offer at most two of the following three at the same time, once a network partition is actually happening:
- Consistency — every read returns the latest write, or an error. This is strong consistency, specifically framed for CAP.
- Availability — every request to a node that's still up gets a real, non-error answer — not necessarily the newest value, just a valid one.
- Partition tolerance — the system keeps functioning even when messages between nodes get dropped or delayed.
People commonly misread this as "pick two of the three, permanently, and design around that choice forever." Real networks partition all the time, though — a cable gets severed, a switch dies mid-afternoon — so partition tolerance was never really a box you get to leave unchecked; it's just what running on an imperfect network means. The actual decision CAP describes only ever comes up while a partition is happening: when one part of Fanline's system genuinely can't reach another part, does it keep responding anyway, even if that response might be stale (that's AP), or does it stop responding until it can be sure the answer is correct (that's CP)? The moment the network is behaving normally again, there's no contradiction at all in offering both at once — CAP is silent on the ordinary case, and for Fanline the ordinary case is nearly every second of the day.
| CP path | AP path | |
|---|---|---|
| Under a partition | Declines requests it can't verify are correct | Keeps answering, possibly with a stale value |
| Fanline example | The seat-reservation write — a wrong answer here means two people holding the same seat | The browse-page seat-count badge — a wrong answer just means a slightly off number for a moment |
| What goes wrong if you pick the other one | A fan is told a seat is theirs and it isn't | A fan sees a seat listed as open that's already sold |
PACELC: What CAP Leaves Out
CAP only covers the rare moment a partition is actually happening; the rest of the time — which for Fanline is nearly every second of every day — it has nothing to say at all. PACELC fills that in: during a Partition, pick Availability or Consistency; Else, pick Latency or Consistency.
That second half is the part that actually matters for a ticketing platform. Even with no partition anywhere near the system, forcing every reservation to coordinate with the leader — no exceptions — adds real latency to the single most time-pressured click on the entire site: the one where a fan is racing a countdown clock against a hundred thousand other fans chasing the same handful of seats. Fanline accepts that latency cost on the reserve path without hesitation, because being forty milliseconds slower beats being wrong about who holds a seat. The browse badge makes the opposite call for the opposite reason: it chases low latency because being off by a seat or two for a moment costs nothing at all.