Skip to main content
CodeOath
← All posts

System Design103 min total · 14 parts

System Design Fundamentals for Interviews: Scalability, Trade-offs, and the Framework Interviewers Actually Grade

Part 3 of 14 · ~5 min

Scalability Fundamentals

Fanline runs on one application server and one database for most of its first year, and honestly, that's the right call — The Locksmith seats 300 people, its shows go on sale a few weeks ahead of time, and even a rough night is a couple hundred requests spread across an hour. "Scalability" gets thrown around loosely in everyday conversation, but here it means something specific: a system's capacity to absorb growth — more users, more data, more traffic — by adding resources rather than rebuilding from scratch. There are only two directions to grow in, and which one you reach for first ends up shaping almost everything that comes after it.

Vertical Scaling

Making one machine bigger — more cores, more memory, a faster disk — is vertical scaling, and it's the path of least resistance because nothing about the application changes. The exact code that ran comfortably on last month's box runs, untouched, on this month's bigger one. When Fanline signs a second venue, a 900-seat hall a few towns over, bumping the app server's specs is a five-minute console change that buys another six months without anyone opening an editor.

But the ceiling arrives sooner than the invoice suggests. There's a largest machine available at any given moment, and well before you reach it, price stops tracking linearly with capacity — a box twice as powerful as the mid-tier option routinely costs far more than double, because the top of any hardware lineup carries its own premium. And no matter how much RAM sits inside it, one machine is still one machine: it goes down, and there's nothing behind it to catch the fall.

Horizontal Scaling

Horizontal scaling takes the opposite bet — add more machines and split the work across them, instead of asking any single one to carry more. There's effectively no ceiling here; need more room, plug in another box. It also buys something vertical scaling structurally can't: lose one server out of a fleet of five, and the other four keep the lights on, slower but alive, rather than the entire platform vanishing along with the one box that used to be everything.

None of that comes for free just because it sounds better on paper. It needs something in front of the fleet deciding where each request goes — load balancing, next chapter — and, more importantly, it needs every server behind that load balancer to be genuinely swappable for any other.

Why Horizontal Scaling Depends on Statelessness

Here's the piece that ties this whole chapter together, and it's easy to skim past without really landing: horizontal scaling only functions because a stateless server can take any request at all and answer it correctly, since it isn't privately holding onto something the rest of the fleet doesn't know about.

Fanline finds out the hard way what happens when that assumption quietly breaks. Traffic starts creeping past what one box handles comfortably, and the obvious move is a second app server behind a load balancer. It ships on a Thursday. By Saturday, support has an irate email: a customer added three seats to her cart for a weekend show, stepped away to look up the venue's address, came back, and her cart was simply empty. Nothing had actually been removed. Her first request had landed on Server A, whose in-memory cart storage was the fastest thing to build under a deadline; her second request landed on Server B, which had never heard of her or her seats. The held tickets were sitting fine in Server A's memory the entire time — but from where she stood, her cart had vanished, and she bought from a competitor before anyone traced the bug.

That's a stateful design quietly breaking its own contract. A customer's follow-up request has to land back on the exact server that handled her first one, or whatever it was holding for her is nowhere to be found. Left uncorrected, the load balancer can't route freely anymore — it has to start pinning each customer to a specific server (session affinity, or "sticky sessions"), and the instant that server dies, every cart pinned to it dies with it. That's the single-point-of-failure problem from vertical scaling again, just now hiding behind a fleet of machines instead of one.

The fix Fanline actually ships is to stop keeping that cart data on the app server at all, and move it into a store every instance can reach the same way:

Before (fragile):
  Customer -> Load Balancer -> Server A (cart lives only in its own memory)
  Her next click MUST land back on Server A, or the cart is simply gone.

After (scalable):
  Customer -> Load Balancer -> any Server (A, B, C — genuinely doesn't matter)
                                    |
                                    v
                          Shared cart store (Redis)
  Every server can answer every request, because no single one of them
  is the only place holding what she put in her cart.

With cart data living in Redis instead of any one server's memory, the fleet becomes fully interchangeable, the load balancer can send traffic to whichever box is least busy, and losing a server costs only whatever request was mid-flight on it — nothing that customer had done is lost, because no single machine was ever the only record of it.

Common mistake: calling a fleet "stateless" while it quietly keeps something that matters in local memory anyway — a per-process cache of a hot event's remaining seats, a short queue of confirmation emails waiting to send. It looks fine on the architecture diagram right up until two requests from the same person land on two different boxes, which is exactly the bug that cost Fanline a sale before anyone found it.