Skip to main content
CodeOath
← All posts

System Design103 min total · 14 parts

System Design Fundamentals for Interviews: Scalability, Trade-offs, and the Framework Interviewers Actually Grade

Part 9 of 14 · ~7 min

API Design Essentials

Whatever contract Fanline's backend signs with a client — the API — ends up dictating things far beyond what data comes back: how quickly it responds, how gracefully it can change later without breaking someone, and whether hitting it twice by accident is safe or catastrophic.

REST, RPC, and GraphQL Side by Side

RESTRPC (e.g. gRPC)GraphQL
Mental modelResources addressed by URL, manipulated with HTTP verbsRemote function calls against a shared schemaOne endpoint; the client names exactly the fields it wants
Over/under-fetchingCommon — a fixed response shape whether or not the client needs all of itMostly a non-issue — each call returns precisely what that method definesDesigned away — the client only ever gets what it asked for
Where Fanline uses itThe public partner API, so a venue's own site can embed a "buy tickets" widget against a stable contractBooking service talking to the inventory service internally, where every millisecond counts and Fanline owns both sidesThe mobile app's home feed, which needs an event plus its venue plus pricing plus waitlist status in one round trip
CostSimple, well understood, but the response shape is rigidNeeds a shared schema (Protocol Buffers) and doesn't play naturally with browsersA single heavy query can strain the server that has to resolve it, and it doesn't cache as cleanly as a plain URL does

None of the three beats the others in every situation — REST's simplicity and per-URL cacheability suit a public partner integration well; gRPC's compact binary format suits internal calls between two services Fanline already owns end to end; GraphQL's client-driven shape suits a mobile screen pulling together several backend resources at once, especially over a connection where every extra round trip is felt.

Pagination: Offset vs. Cursor

  • Offset pagination (?page=3&limit=20) is simple and lets a client jump straight to an arbitrary page. It falls apart the moment rows get inserted or removed mid-scroll: Fanline's "shows near you" feed adds new events constantly, and a fan paging through results while a fresh event lands on page one can see the same show twice, or miss one entirely, as everything shifts under them.
  • Cursor pagination hands back an opaque token alongside each page — typically built from whatever the last row sorted on — and the client feeds that token back to ask for everything past that exact row, instead of asking for a numbered position. Inserts happening around it don't break anything, because "past this specific row" never shifts the way a raw offset would — what it costs you is the ability to jump straight to page seven; getting there means walking every cursor that comes before it first.

Fanline's infinite-scroll feed moved to cursor pagination for that exact reason — a busy, constantly-changing feed is where offset pagination's failure shows up most often, which is precisely when the feed has the most eyes on it.

Idempotency Keys

A plain POST isn't naturally safe to repeat — sending "charge this card $180" twice is supposed to charge it twice, that's the entire point of a create request. That turns into a real problem the moment a connection hiccups: a customer taps "confirm purchase" on shaky signal, the charge goes through on Fanline's end, but the success response never makes it back to her phone. From where she's sitting, that looks identical to nothing happening at all, and tapping the button again is the obvious next move.

That's exactly what happened once, before idempotency keys existed on this endpoint: a customer on a train with patchy coverage got billed twice for the same three tickets, because her app's retry logic read the missing response as a failure and resubmitted the request. An idempotency key closes that gap — the client stamps a unique key onto a given purchase attempt, and the server remembers which keys it's already handled. Show up with the same key twice, and the server hands back the original response instead of charging the card again:

POST /checkout
Idempotency-Key: 7c2a9e10-4b3f-4a8e-9c11-...
{ "eventId": "marlow-vance-msg", "seats": ["F12", "F13"], "amount": 180 }

Server logic:
  if this idempotency key has been seen before:
      hand back the response stored from that first time
  else:
      charge the card, create the order
      remember (idempotency_key -> response)
      return the response

It's the identical underlying problem as at-least-once queue delivery, just dressed up as a synchronous API instead of a queue consumer — a duplicate attempt at the same logical action is the failure mode either way, and turning a repeat into a harmless no-op is the fix either way too.

Rate-Limiting Algorithms

Rate limiting puts a ceiling on how many requests one client can make in a given stretch of time, and it stops being a nice-to-have for Fanline the moment scripts start hammering the reserve-seat endpoint faster than any human ever could, trying to lock up whole blocks of a hot show's inventory to resell.

Token bucket fills a bucket to capacity tokens over time at a fixed rate; each request spends one, and an empty bucket means the request gets turned away. Short bursts up to the bucket's size sail through while the average rate over time still holds:

function tryConsumeToken(clientBucket):
    timeSinceRefill = currentTime() - clientBucket.lastRefill
    refilled = timeSinceRefill * clientBucket.refillRate
    clientBucket.tokens = min(clientBucket.capacity, clientBucket.tokens + refilled)
    clientBucket.lastRefill = currentTime()

    if clientBucket.tokens < 1:
        return "throttled"
    clientBucket.tokens -= 1
    return "let it through"

Leaky bucket queues incoming requests and drains them at a strictly fixed rate no matter how bursty the arrivals were — the opposite instinct from token bucket, smoothing everything into an even stream instead of letting a burst straight through, at the cost of extra wait time for whoever's queued during that burst.

Fixed window counters simply count requests inside a fixed slice of time and zero out at each boundary. Cheap to build, but it has an edge-burst hole: a client can spend its full allowance right before a window ends and its full allowance again right after the next one starts, nearly doubling the intended rate across that seam — exactly the kind of gap a determined scalping script finds within its first hour of trying.

Sliding window counters skip the hard reset entirely — they blend in a fraction of the previous window's count, scaled by how much that previous window still overlaps the current moment, which produces a smoothed-out estimate without the overhead of logging every single request's exact timestamp:

function allowRequest(currentCount, previousCount, elapsedIntoWindow, windowSize, limit):
    weight = (windowSize - elapsedIntoWindow) / windowSize
    estimatedCount = previousCount * weight + currentCount
    if estimatedCount < limit:
        return True
    return False
AlgorithmHandles burstsEvens out the rateHow hard to build
Token bucketYes, up to bucket sizeNo — a burst passes straight throughLow
Leaky bucketNo — queued, drained at a constant paceYesModerate
Fixed windowYes — and can double up across a boundaryNoSimplest
Sliding windowPartially, smoothed outMostlyMedium

Fanline settles on token bucket for the reserve-seat endpoint specifically because genuine fans are bursty too — someone mashing refresh in the opening seconds of an on-sale looks, in isolation, a lot like a bot, and leaky bucket's constant drip would have throttled that real burst exactly as hard as a scripted one. Allowing a short legitimate burst while still capping the sustained rate no real person keeps up for more than a few seconds turns out to separate the two crowds well enough to matter. There's a question this section has been dodging on purpose, though — once more than one app server is watching the same client, whose bucket is actually the real one? The second worked example below is where that gets answered.