System Design103 min total · 14 parts
System Design Fundamentals for Interviews: Scalability, Trade-offs, and the Framework Interviewers Actually Grade
Part 4 of 14 · ~5 min
Load Balancing
Something has to sit in front of Fanline's app-server fleet and decide, for each request that arrives, which box actually handles it. The algorithm behind that decision matters more than it looks, because picking wrong either leaves half the fleet idle or funnels traffic straight at the servers least able to take more of it.
Common Algorithms
| Algorithm | Decision rule | Where it shines | Where it falls short |
|---|---|---|---|
| Round robin | Cycle through the fleet in order, one request per server | Every request costs roughly the same to serve | Blind to load — a box wrestling with one slow checkout still gets handed the next request |
| Least connections | Hand the request off to whichever box has the fewest open connections right now | Request cost varies wildly (a quick browse next to a full checkout) | A little extra bookkeeping versus round robin, usually a fair trade |
| Weighted variants | Same idea, but bigger boxes earn a proportionally bigger share | A fleet with mismatched hardware | Weights drift out of date whenever the hardware mix changes |
| Consistent hashing | Place servers and keys on a ring; a key belongs to the nearest server clockwise from it | The same key has to reach the same server, request after request | Genuinely trickier to build correctly than the others |
For anything where one server is as good as another, round robin or least connections do the job without drama. Consistent hashing solves a completely different kind of problem — it's for when a repeated request for the same key needs to keep finding the same server, usually because that server is the one holding something useful for that key in memory.
Consistent Hashing: What It's Actually Fixing
Fanline runs headfirst into this the day it starts keeping a hot show's seat map cached in each app server's own memory, purely to skip a database round trip on every seat click during a rush. With plain round robin sitting in front, two clicks from the same person on the same seat map land on two different servers roughly four times out of five — so the in-memory cache barely helps; it's practically a coin toss whether the server answering this click already has the map warmed up.
The obvious first attempt — route by hash(event_id) % number_of_servers — works fine until the fleet size changes. Add a sixth box for the next big on-sale, and that modulo's denominator changes, which flips the result for nearly every event, not only the handful that genuinely needed to move. Every cache that was warm a second ago is now serving the wrong show, which looks exactly like wiping the cache clean at the precise moment you added capacity to relieve pressure on it.
Consistent hashing sidesteps this by placing servers on positions around a fixed circle (commonly spanning 0 through 2³²−1, wrapping back to zero) and placing each event on that same circle. An event belongs to whichever server sits first when you walk clockwise from the event's own position:
Positions on the ring climb as you move clockwise, looping back to zero at the top:
Server 1 (pos 15)
/ \
Server 4 (pos 310) Server 2 (pos 95)
\ /
Server 3 (pos 210)
"marlow-vance-msg" hashes to 40 -> Server 2 owns it (first clockwise from 40)
"locksmith-jazznight" hashes to 260 -> Server 4 owns it (first clockwise from 260)
Drop a sixth server onto position 65, and the only slice that changes hands is the one between the previous server clockwise and that new position — positions 16 through 65, previously Server 2's, now belong to the newcomer instead. Every other slice on the ring stays exactly as it was: Server 3, Server 4, and most of what Servers 1 and 2 already held remain untouched and still warm. Take a server out on a quiet Tuesday instead, and it plays out in reverse — its slice falls to whoever sits next clockwise, and nothing further around the ring is disturbed at all.
Boiled down: resizing the fleet stops being an event that scrambles nearly every cache entry, and becomes one that only touches the slice that belonged to whichever server just joined or left. Real deployments also give each physical server several positions on the ring ("virtual nodes") instead of one, purely so that a single unlucky box doesn't end up owning a disproportionately wide slice by pure chance — without that spread, Fanline's single busiest machine could easily be the one that happened to land the widest arc.
L4 vs. L7 Load Balancing
There's a second axis entirely — which layer the balancer is actually reading to make its call — and it's what determines whether the trick above is even possible.
- Layer 4 balancers only ever see an IP address and a port. A connection gets forwarded without anyone peeking inside it, which is why L4 is quick and uncomplicated — and also why it has zero opinion on what the request is actually asking for.
- Layer 7 balancers understand the protocol riding on top — HTTP, in this case — and can make a decision based on a URL path, a header, a cookie, or even the body. Pulling an event ID out of
GET /events/marlow-vance-msg/seatsso it can be hashed only works at this layer; an L4 balancer is working one level below where a URL even exists, so it literally cannot see it.
Cost favors L4, since there's simply less for it to examine per packet. Fanline needs L7 the moment a routing decision depends on anything beyond which IP hit which port — and once seat-map caching enters the picture, that's true immediately.