Skip to main content
CodeOath
← All posts

Architecture & Patterns68 min total · 17 parts

Microservices vs. Monolith: The Trade You're Actually Making

Part 10 of 17 · ~3 min

Resilience Patterns

The network hop that replaced an in-process call which basically never failed — flagged a few chapters back — needs deliberate handling now, or one struggling service quietly drags every one of its callers down with it. Furrow's worst version of this happens on a Wednesday morning: forty-some farms are all uploading their weekly harvest counts inside the same ninety-minute window, Harvest's database is under real write pressure from the bulk imports, and every checkout request Subscriptions sends it starts taking four, five, six seconds to answer instead of the usual fifty milliseconds.

  • Timeouts. Subscriptions' calls to Harvest had no timeout at all in the first version — a genuinely slow-but-not-down Harvest instance just held the caller's connection open, and enough of those piling up during the Wednesday rush exhausted Subscriptions' own connection pool, which meant checkout started failing for everyone, including subscribers whose zone had nothing to do with the slow farms. A four-second timeout, tuned to what a healthy Harvest call actually takes, turned "checkout is completely down" into "a few checkouts fail and can retry" — a much smaller, much more honest failure.
  • Retries, with backoff and jitter. A single slow response is often just that — a momentary blip, worth one retry. But have every one of Subscriptions' instances retry right away, and retry hard, and a Harvest instance that was merely struggling can get tipped straight into fully overloaded — a pile-up made entirely of good intentions. Exponential backoff spaces retries out further each time, which helps, but it isn't the whole fix on its own: if every instance backs off on the exact same schedule, they end up retrying in near-perfect sync anyway, hitting Harvest in synchronized waves instead of one steady trickle. Adding jitter — a small random offset on top of the backoff delay — is what actually breaks that synchronization and spreads the retries out over time instead of bunching them back up.
  • Idempotency, for the calls that get retried. A retried "reserve one unit of lettuce" call is only safe to repeat if reserving it twice by accident doesn't reserve two units — Harvest's reservation endpoint takes a client-generated request ID and treats a repeat of the same ID as "already done, here's the original result" rather than doing the work again. Without that, the timeout-and-retry logic above quietly turns "the first request probably succeeded, we just didn't hear back in time" into a second, real reservation nobody meant to make.
  • Circuit breakers. Once calls to Harvest have failed enough times in a row, Subscriptions' circuit breaker trips into an open state and simply refuses to place the call for a short cooldown, returning an instant failure instead of making every caller sit through its own timeout — which gives Harvest breathing room to work through its backlog instead of getting buried under retry traffic on top of it. Coming back isn't all-or-nothing: the breaker first goes half-open, letting exactly one probe request through to check whether Harvest has genuinely recovered before it commits to closing fully or falling back to failing fast.
  • Bulkheads. Named for the sealed compartments in a ship's hull. Furrow keeps the connection pool checkout uses to call Harvest completely separate from the one the nightly reconciliation job uses for the same service — so if the reconciliation job's queries back up, it can't exhaust the pool checkout depends on, and a batch job running slow on a Tuesday night never once threatens a Wednesday-afternoon customer trying to place an order.

These stack up in practice rather than standing alone: a well-behaved call from Subscriptions to Harvest gives up after a set timeout, retries only a capped number of times with backoff and jitter spacing them out, sits behind a breaker that can cut it off entirely, tolerates being repeated safely if it does retry, and draws from a connection pool that's walled off from everything else Subscriptions does — five distinct, individually unremarkable protections, each one covering a failure mode none of the other four touch.