Skip to main content
CodeOath
← All posts

System Design103 min total · 14 parts

System Design Fundamentals for Interviews: Scalability, Trade-offs, and the Framework Interviewers Actually Grade

Part 8 of 14 · ~5 min

Message Queues and Asynchronous Processing

Everything covered until now has quietly assumed a caller sticks around waiting for its answer. A message queue exists to break that assumption on purpose: whatever produces a message hands it off and gets straight back to what it was doing, and whatever eventually consumes it does so on its own clock, entirely unlinked from the producer's timing — or even from whether the producer is still around.

Two Ways a Queue Can Fan a Message Out

  • Point-to-point picks a single winner from a pool of consumers for each message, which is mainly how work gets spread across a group of workers. Fanline's e-ticket PDFs render this way: a pool of ticket-renderer workers all pull from the same queue of completed orders, and each order's PDF gets built exactly once, by whichever worker grabs it first.
  • Pub/sub hands each message to every subscriber listening on that topic. A "ticket purchased" event needs to reach the confirmation-email service, the analytics pipeline, and — when it applies — the waitlist service all at once; each needs its own independent copy of the same event, which a point-to-point queue simply can't hand out.

When a Queue Is Earning Its Keep — and When It's Just Extra Moving Parts

The team's first version of checkout does everything in one synchronous pass: charge the card, write the order, render the ticket PDF, send the confirmation email, then finally tell the customer it worked — all inside one HTTP request. It's fine until the email provider has a slow afternoon, and suddenly every "buy" button on the site hangs for six or seven seconds waiting on a mail server that has nothing to do with whether the seat is actually theirs.

A queue earns its place once at least one of a few things is true: producer and consumer genuinely need to scale on different curves (a burst of purchases shouldn't force the email service to grow in lockstep — it can just work through a backlog), the consumer being briefly unavailable shouldn't take the producer down with it (exactly what decoupled checkout from that flaky mail provider), or the work has to fan out to several independent listeners reacting to one event, the way "ticket purchased" does. Fanline's fix keeps the checkout transaction itself synchronous — charge, write the order, done — and pushes everything downstream of "the order exists" onto a queue as one published event, with email, rendering, and analytics each subscribing on their own. Checkout latency drops back to whatever the payment gateway itself takes, since nothing else is in the critical path anymore.

It would've been the wrong call to put a queue on the one part of checkout that genuinely has to stay synchronous, though: the "is this seat still free" check. A queue can't hand back an instant yes or no — the entire design of a queue is fire-and-forget — and a customer staring at the seat map needs an answer right now, not eventually.

Common mistake: wedging a queue between two services purely on the theory that "that's how scalable systems are built," and only later realizing the caller genuinely needed an answer right away and is now stuck awkwardly polling for one — which brings back the exact delay the queue was meant to erase, just with extra moving parts along the way.

At-Least-Once vs. Exactly-Once Delivery (and the Dead Letters They Leave Behind)

Most real-world queues offer at-least-once delivery as their guarantee: a message is going to arrive, but nothing stops it from arriving more than once — if a consumer dies in the gap between finishing the work and confirming it, that same message shows back up once the consumer is healthy again. Which means consumers need to be idempotent by design: running the same message through twice has to leave things exactly where running it once would have, or every duplicate delivery becomes a duplicate side effect in the real world. Fanline ran into this directly — the confirmation-email consumer crashed mid-send during a deploy, and the redelivered "ticket purchased" event sent one customer two identical confirmation emails for the same order. No card was charged twice, since the charge lived in the synchronous checkout path rather than the queue, but the fix followed the same pattern regardless: the consumer now checks whether it's already sent a confirmation for that order ID before doing anything, so a redelivery quietly becomes a no-op instead of a repeat.

Exactly-once delivery promises something that sounds strictly better, but actually pulling it off over a network that drops and reorders things is a much harder engineering problem than the marketing implies. Look under the hood of most systems that claim it and you'll usually find at-least-once delivery with a deduplication layer bolted onto the consumer side — which is the identical idempotency requirement, just wearing a different name tag. Making your own consumers idempotent as a baseline habit beats trusting any vendor's "exactly-once" label to mean your own code is off the hook.

There's a failure neither guarantee covers on its own: a message that's structurally broken — an out-of-date mobile client sending a "ticket purchased" event in a schema the renderer no longer recognizes — gets redelivered, fails, gets redelivered again, forever, sitting at the head of the queue and blocking everything stacked up behind it. A dead-letter queue is the fix: after a capped number of failed attempts, the message gets pulled off the main queue and dropped into a separate one for a human to look at, rather than retried without end. Fanline didn't have one for the queue's first month, and paid for it when a batch of malformed events from a since-patched app bug sat retrying at the front of the ticket-render queue for the better part of an hour, quietly stalling every legitimate order stacked up behind them.