System Design103 min total · 14 parts
System Design Fundamentals for Interviews: Scalability, Trade-offs, and the Framework Interviewers Actually Grade
Part 9 of 14 · ~7 min
API Design Essentials
Whatever contract Fanline's backend signs with a client — the API — ends up dictating things far beyond what data comes back: how quickly it responds, how gracefully it can change later without breaking someone, and whether hitting it twice by accident is safe or catastrophic.
REST, RPC, and GraphQL Side by Side
| REST | RPC (e.g. gRPC) | GraphQL | |
|---|---|---|---|
| Mental model | Resources addressed by URL, manipulated with HTTP verbs | Remote function calls against a shared schema | One endpoint; the client names exactly the fields it wants |
| Over/under-fetching | Common — a fixed response shape whether or not the client needs all of it | Mostly a non-issue — each call returns precisely what that method defines | Designed away — the client only ever gets what it asked for |
| Where Fanline uses it | The public partner API, so a venue's own site can embed a "buy tickets" widget against a stable contract | Booking service talking to the inventory service internally, where every millisecond counts and Fanline owns both sides | The mobile app's home feed, which needs an event plus its venue plus pricing plus waitlist status in one round trip |
| Cost | Simple, well understood, but the response shape is rigid | Needs a shared schema (Protocol Buffers) and doesn't play naturally with browsers | A single heavy query can strain the server that has to resolve it, and it doesn't cache as cleanly as a plain URL does |
None of the three beats the others in every situation — REST's simplicity and per-URL cacheability suit a public partner integration well; gRPC's compact binary format suits internal calls between two services Fanline already owns end to end; GraphQL's client-driven shape suits a mobile screen pulling together several backend resources at once, especially over a connection where every extra round trip is felt.
Pagination: Offset vs. Cursor
- Offset pagination (
?page=3&limit=20) is simple and lets a client jump straight to an arbitrary page. It falls apart the moment rows get inserted or removed mid-scroll: Fanline's "shows near you" feed adds new events constantly, and a fan paging through results while a fresh event lands on page one can see the same show twice, or miss one entirely, as everything shifts under them. - Cursor pagination hands back an opaque token alongside each page — typically built from whatever the last row sorted on — and the client feeds that token back to ask for everything past that exact row, instead of asking for a numbered position. Inserts happening around it don't break anything, because "past this specific row" never shifts the way a raw offset would — what it costs you is the ability to jump straight to page seven; getting there means walking every cursor that comes before it first.
Fanline's infinite-scroll feed moved to cursor pagination for that exact reason — a busy, constantly-changing feed is where offset pagination's failure shows up most often, which is precisely when the feed has the most eyes on it.
Idempotency Keys
A plain POST isn't naturally safe to repeat — sending "charge this card $180" twice is supposed to charge it twice, that's the entire point of a create request. That turns into a real problem the moment a connection hiccups: a customer taps "confirm purchase" on shaky signal, the charge goes through on Fanline's end, but the success response never makes it back to her phone. From where she's sitting, that looks identical to nothing happening at all, and tapping the button again is the obvious next move.
That's exactly what happened once, before idempotency keys existed on this endpoint: a customer on a train with patchy coverage got billed twice for the same three tickets, because her app's retry logic read the missing response as a failure and resubmitted the request. An idempotency key closes that gap — the client stamps a unique key onto a given purchase attempt, and the server remembers which keys it's already handled. Show up with the same key twice, and the server hands back the original response instead of charging the card again:
POST /checkout
Idempotency-Key: 7c2a9e10-4b3f-4a8e-9c11-...
{ "eventId": "marlow-vance-msg", "seats": ["F12", "F13"], "amount": 180 }
Server logic:
if this idempotency key has been seen before:
hand back the response stored from that first time
else:
charge the card, create the order
remember (idempotency_key -> response)
return the response
It's the identical underlying problem as at-least-once queue delivery, just dressed up as a synchronous API instead of a queue consumer — a duplicate attempt at the same logical action is the failure mode either way, and turning a repeat into a harmless no-op is the fix either way too.
Rate-Limiting Algorithms
Rate limiting puts a ceiling on how many requests one client can make in a given stretch of time, and it stops being a nice-to-have for Fanline the moment scripts start hammering the reserve-seat endpoint faster than any human ever could, trying to lock up whole blocks of a hot show's inventory to resell.
Token bucket fills a bucket to capacity tokens over time at a fixed rate; each request spends one, and an empty bucket means the request gets turned away. Short bursts up to the bucket's size sail through while the average rate over time still holds:
function tryConsumeToken(clientBucket):
timeSinceRefill = currentTime() - clientBucket.lastRefill
refilled = timeSinceRefill * clientBucket.refillRate
clientBucket.tokens = min(clientBucket.capacity, clientBucket.tokens + refilled)
clientBucket.lastRefill = currentTime()
if clientBucket.tokens < 1:
return "throttled"
clientBucket.tokens -= 1
return "let it through"
Leaky bucket queues incoming requests and drains them at a strictly fixed rate no matter how bursty the arrivals were — the opposite instinct from token bucket, smoothing everything into an even stream instead of letting a burst straight through, at the cost of extra wait time for whoever's queued during that burst.
Fixed window counters simply count requests inside a fixed slice of time and zero out at each boundary. Cheap to build, but it has an edge-burst hole: a client can spend its full allowance right before a window ends and its full allowance again right after the next one starts, nearly doubling the intended rate across that seam — exactly the kind of gap a determined scalping script finds within its first hour of trying.
Sliding window counters skip the hard reset entirely — they blend in a fraction of the previous window's count, scaled by how much that previous window still overlaps the current moment, which produces a smoothed-out estimate without the overhead of logging every single request's exact timestamp:
function allowRequest(currentCount, previousCount, elapsedIntoWindow, windowSize, limit):
weight = (windowSize - elapsedIntoWindow) / windowSize
estimatedCount = previousCount * weight + currentCount
if estimatedCount < limit:
return True
return False
| Algorithm | Handles bursts | Evens out the rate | How hard to build |
|---|---|---|---|
| Token bucket | Yes, up to bucket size | No — a burst passes straight through | Low |
| Leaky bucket | No — queued, drained at a constant pace | Yes | Moderate |
| Fixed window | Yes — and can double up across a boundary | No | Simplest |
| Sliding window | Partially, smoothed out | Mostly | Medium |
Fanline settles on token bucket for the reserve-seat endpoint specifically because genuine fans are bursty too — someone mashing refresh in the opening seconds of an on-sale looks, in isolation, a lot like a bot, and leaky bucket's constant drip would have throttled that real burst exactly as hard as a scripted one. Allowing a short legitimate burst while still capping the sustained rate no real person keeps up for more than a few seconds turns out to separate the two crowds well enough to matter. There's a question this section has been dodging on purpose, though — once more than one app server is watching the same client, whose bucket is actually the real one? The second worked example below is where that gets answered.