System Design103 min total · 14 parts
System Design Fundamentals for Interviews: Scalability, Trade-offs, and the Framework Interviewers Actually Grade
Part 5 of 14 · ~7 min
Caching Strategies
A cache spends a small amount of fast, pricier storage to buy a large cut in latency and load on whatever's sitting behind it — for Fanline, that's nearly always the database holding event details and live seat counts. The whole strategy comes down to two questions: what earns a spot in the cache in the first place, and what gets kicked out once the space is needed for something else?
Cache Population Strategies
| Strategy | Mechanics | What it costs |
|---|---|---|
| Cache-aside | App checks the cache first; on a miss, it reads the database and fills the cache itself | Only ever caches what someone actually asked for, but the first request for any event eats a guaranteed miss |
| Read-through | Same idea, but the cache library owns the miss-and-fill step instead of app code doing it by hand | Less application logic to get wrong, but the cache layer needs a direct line to the database |
| Write-through | Every write lands in the cache and the database together, as one synchronous step | The cache is never behind, but every write now pays for two round trips instead of one |
| Write-back | The write hits the cache immediately; the database catches up later, asynchronously | Fast writes, but anything sitting only in cache disappears if the cache dies first |
Fanline leans on cache-aside for event details and seat maps, for the same reason nearly everyone does: it only burns cache space on shows people are actively looking at — an event three months out that nobody's opened yet never takes up a slot. The pattern reads the same every time it shows up:
function getEventDetails(eventId):
cached = cache.get(eventId)
if cached is not None:
return cached # hit
details = database.getEvent(eventId) # miss — go ask the database, which never lies
cache.set(eventId, details)
return details
Write-through would keep that cache permanently in sync with the database, at the price of every write waiting on both systems — a bad trade on Fanline's checkout path, where every extra millisecond is a millisecond a customer might spend clicking "buy" on a rival tab instead. Write-back flips it around: a write returns quickly since the cache is the only thing that has to confirm it landed, but that leaves a stretch of time where the truest number for "seats remaining" exists nowhere except somewhere volatile — precisely the kind of gap a ticketing platform can't tolerate, for reasons the CAP chapter unpacks properly later on.
TTL, Eviction, and the Stampede Nobody Plans For
Nothing stays in a cache forever, so two separate mechanisms usually cooperate to clear room: a TTL, after which an entry is treated as stale and dropped no matter how popular it's been, and an eviction policy, the rule that picks what actually gets sacrificed once there's no free space left for something new.
| Policy | Removes | Suits | Struggles with |
|---|---|---|---|
| LRU | Whatever hasn't been touched in the longest stretch | General-purpose traffic, where recency is a decent stand-in for what gets asked for next | One unusual wave of scan-like lookups can knock out data that's still genuinely popular |
| LFU | Whatever has the lowest total access count | A stable "hot set" that's accessed far more than the rest of the catalog | A once-popular entry can linger long after it stopped mattering, propped up by an old count |
Fanline runs LRU by default and it's fine — right up until a band that played The Locksmith three years back announces a surprise reunion show. That event's cache entry evicted itself out of every app server months ago, since nobody had looked at it since. The announcement goes out, tens of thousands of fans hit that page inside the same sixty seconds, and every one of those requests misses cache at the same instant, all falling through to the database together — a cache stampede. It's a different failure than an ordinary miss: it's the same miss landing thousands of times in the same breath instead of being spread across a day, and that alone is enough to knock the database over even though the steady-state load for that event, spread out normally, would be nothing.
The fix isn't a bigger cache or a longer TTL — it's making sure exactly one of those thousands of simultaneous misses actually reaches the database, while the rest wait on that one request and share its answer:
function getEventDetails(eventId):
cached = cache.get(eventId)
if cached is not None:
return cached
if not acquireRebuildLock(eventId): # somebody else is already refilling this key
wait briefly, then re-check cache.get(eventId)
return cached
details = database.getEvent(eventId) # only the lock-holder actually reaches the DB
cache.set(eventId, details)
releaseRebuildLock(eventId)
return details
Call it single-flight, or request coalescing — either way, it turns "ten thousand simultaneous database hits" into "one database hit plus ten thousand short waits," which the database survives comfortably where the first version doesn't. It's a sharper version of the same lesson eviction already teaches: a cold-key miss has a cost, and the moment a lot of traffic can converge on the same cold key at once, that cost lands all at once instead of amortizing across the day.
Pushing Static Assets Out to a CDN
A content delivery network scatters caching servers (edge nodes) around the globe, physically near the people actually requesting things, and holds onto whatever's identical for every visitor and doesn't change often — a venue's promo photo, an artist's poster art, the JS and CSS bundles the storefront ships. Someone browsing from Mumbai pulling up artwork for a US tour date shouldn't need a round trip to a server on another continent for that; the nearest edge node already has the image cached and serves it directly, shrinking both the physical distance and the number of repeat hits the origin server ever has to absorb.
That's a strong fit for anything uniform across every viewer. It falls apart for anything personal, or anything that has to reflect whatever was just written a second ago — a customer's own live seat count being the obvious example, for the same underlying reason: a cache only pays for itself when a lot of requests can share a single cached answer, and a number that changes on every completed checkout has nothing to share.
Cache Invalidation: The Genuinely Hard Part
Computer science has an old running joke that only two things are genuinely hard: naming things, and knowing when a cache has gone stale. The second half of that joke holds up under scrutiny. Nothing about a cache is magic — it's just a duplicate answer parked somewhere fast, and duplicates drift out of sync with whatever they're shadowing. The instant the real record changes, the duplicate is simply wrong until some mechanism catches that and reacts to it, and building a mechanism that catches it reliably, under real traffic, is where the difficulty actually lives.
- Let it expire. A TTL lets an entry go stale for a bounded stretch and accepts that cost outright. Cheap to build, but any read inside that window can hand back an old answer, and the TTL itself is really just a dial between how stale and how fast you're willing to be.
- Clear it on write. The instant the underlying row changes, the same write that changed it also clears or refreshes the matching cache entry. Solid, provided every path capable of touching that row remembers to do it — a one-off refund script that skips the normal checkout code, for instance, and the cache quietly falls out of sync with nothing around to notice.
- Clear it on a broadcast event. Rather than trusting every write path to remember by hand, the data layer itself announces that something changed, and any cache that cares subscribes and clears itself independently. More durable than hoping every write path behaves, at the cost of the extra plumbing needed to publish and consume those announcements reliably.
Fanline builds write-triggered invalidation straight into the reservation transaction on purpose: the instant a seat is held, wiping its cached availability happens inside the same transaction as the hold, not tacked on afterward — because a seat-map cache that's wrong in the optimistic direction, still showing a taken seat as free, is a double-booking waiting to happen, which is a far worse outcome than a taken seat showing as unavailable a beat longer than strictly necessary.
Common mistake: stretching out the TTL whenever a stale-cache bug turns up, instead of tracking down whatever's actually broken in the invalidation logic. Adding more time before expiry doesn't correct anything — it just gives the same underlying hole a wider window to go unnoticed in.