Skip to main content
CodeOath
← All posts

Architecture & Patterns68 min total · 17 parts

Microservices vs. Monolith: The Trade You're Actually Making

Part 6 of 17 · ~6 min

Data Ownership and the Saga Pattern

Whatever a service owns, it owns alone — nobody else reaches in and touches it. This is the single rule that makes the rest of the independence real, and it's also the one Furrow almost broke without noticing. Early on, someone on the Fulfillment team — under deadline pressure, reasonably — wired the new Subscriptions service to read available_quantity straight out of Harvest's Postgres database, instead of calling Harvest's API for it, because it was faster to ship that week. It worked, right up until Harvest changed that column's meaning during a schema migration and Subscriptions silently started showing wrong availability numbers for two days before anyone noticed. A schema change that should have been Harvest's business alone had quietly become something two teams had to coordinate — which is exactly the coordination cost splitting the services was supposed to remove.

Once ownership is genuinely respected, though, you lose something real: the monolith's cross-table transaction. The old version of "confirm this box" ran as one Postgres transaction:

BEGIN;
UPDATE harvest_lots SET reserved_units = reserved_units + 1 WHERE lot_id = 'lettuce-z3-wk32';
UPDATE subscriber_accounts SET balance_cents = balance_cents - 2400 WHERE id = 'sub_88214';
COMMIT;

Both rows change or neither does — Postgres guarantees it. Split across Harvest's database and Billing's database, there is no single transaction that can span both anymore. The Saga pattern swaps that one guarantee for a chain of smaller, per-service transactions, each one paired with a compensating action that can walk it back if something further down the chain goes wrong:

1. Subscriptions: create the box order, status "pending"
2. Harvest: reserve the produce for the order
   -> if this step fails: Subscriptions cancels the order
3. Billing: charge the subscriber's card
   -> if this step fails: Harvest releases the reservation it made in step 2
                           Subscriptions cancels the order
4. Subscriptions: mark the order confirmed

Sit with what this does and doesn't guarantee. Between the moment Harvest reserves the produce and the moment Billing's charge actually clears, there's a real stretch of time where a specific head of lettuce belongs to a subscriber who hasn't paid for it yet — no snapshot taken during that stretch is "correct" the way a database transaction promises correctness. The system only becomes consistent once every step in the sequence has run, forward or compensating, which is what eventual consistency actually means here: not that correctness is optional, but that it arrives after a sequence finishes rather than all at once. And "release the reservation" isn't a true rollback — if the card charge in step 3 fails after the produce was genuinely marked reserved for a few hundred milliseconds, the compensating action doesn't erase that moment, it just frees the units back to the pool from wherever they currently sit. Designing a saga well means designing every step so it can actually be compensated — Harvest's API needed a releaseReservation endpoint from day one, not bolted on after the first time step 3 failed in production.

Who decides the next step: orchestration vs. choreography

There are two honest ways to wire a saga together, and Furrow's first attempt picked the wrong one for the wrong reason. Choreography means each service reacts to events from the one before it — Harvest publishes "reserved," Billing is listening for that event and charges the card on its own, Subscriptions is listening for the charge result and confirms or cancels. Nobody is in charge; the sequence emerges from everyone knowing which events to react to. It's appealing because no single service has to know about the whole flow. It's also how Furrow lost a full afternoon tracing a stuck order: a deploy briefly took Billing's event listener offline, and there was no single piece of code anywhere that knew "step 3 should have happened by now and didn't" — just a reservation sitting there, correctly made, with nothing watching for its follow-up to go missing.

Orchestration puts one component — a saga coordinator living in Subscriptions — in charge of calling each step in order and deciding what happens on failure. It's more code, and it puts a specific service in the position of knowing about the whole flow, which is a little bit of the coupling the split was supposed to remove. But it also means the sequence lives in exactly one readable place, and a stuck order becomes "the coordinator is waiting on step 3" instead of a mystery. Furrow moved the box-confirmation saga to orchestration after the stuck-order afternoon, and kept choreography for genuinely independent reactions — the payout calculation and the SMS notification, which don't need to report back to anyone and were never the problem.

The gap choreography exposed: writing the event and the database row atomically

There's a smaller trap sitting inside all of this that's easy to miss even once you've picked orchestration: Harvest's reservation step both writes a row to its own database and has to tell the rest of the system it happened, by publishing an event or letting the orchestrator know. Those are two separate operations, and if the process crashes between them — the database commit lands, but the message never gets published — the reservation is real and nobody downstream ever finds out. Furrow hit this exactly once, in testing, when a pod got killed by a rolling deploy in the half-second gap between the two calls, and a reserved lot sat invisible to Billing until someone found it during a weekly reconciliation job.

The fix is the outbox pattern: instead of writing the row and publishing the event as two separate operations, Harvest writes the reservation row and a row describing the event, in the same local database transaction — so either both happen or neither does, because it's one transaction against one database. A small separate process then reads new rows out of that outbox table and actually publishes them, retrying until it succeeds, and marks them published once it has. The event can end up published more than once if that background process crashes mid-retry, which is why every consumer of it — Billing included — still has to treat "confirm this order again" as a safe no-op rather than assuming it will only ever hear about a reservation exactly once.

Common mistake: treating a Saga's eventual consistency as though it were the same guarantee as the transaction it replaced, and skipping the design work for what a mid-sequence failure actually does to a customer. A step that can't be compensated isn't a smaller version of the problem — it's the saga not actually working, discovered in production instead of in a design review.