Skip to main content
CodeOath
← All posts

CI/CD & DevOps65 min total · 17 parts

CI/CD Pipelines Explained: From Push to Production

Part 15 of 17 · ~2 min

Rollbacks

Two weeks after worker starts running three instances, a genuinely bad deploy goes out: a change meant to store each receipt's detected currency adds a new currency_code column to the receipts table, NOT NULL, no default, as part of the same release that also updates worker to populate it. The migration runs first, the way migrations always do, and for about ninety seconds — while the rolling update is still swapping instances — the two old worker instances that haven't been replaced yet try to insert a receipt row the ordinary way, without a currency_code, and the database rejects every single one of them.

The instinct in the moment is the obvious one: roll the code back to the previous version, the one that doesn't know about currency_code at all, and buy time to fix it properly. That instinct is exactly backwards, and it's the surprise worth sitting with: the column is still there, still NOT NULL, still with no default, because rolling back a deploy rolls back code, not the database schema that deploy's migration already changed. The "rolled back" old worker code doesn't set currency_code either — it has never heard of the column — so every insert keeps failing, identically, after the rollback as before it. The team spends several confused minutes staring at a dashboard that looks unchanged, wondering why undoing the bad deploy didn't undo the bad symptom, before someone actually reads the error and realizes the schema was never part of what "rolling back" touched.

The real fix — and the thing that should have shipped in the first place — is staging the change so every step of it is safe for both an old and a new worker to be running against at once, which is exactly the constraint the rolling-updates chapter flagged and nobody connected to a schema change until it actually broke something:

  1. Add currency_code as nullable, no default required. Both old and new code can insert rows fine; old code just leaves it empty.
  2. Deploy worker code that writes currency_code going forward. Old rows still have it empty, which is fine — nothing's reading it yet.
  3. Backfill the empty rows separately, on its own schedule, with no deploy involved.
  4. Only in a later, separate release — once nothing in production could possibly still be writing rows without it — make the column NOT NULL.

None of that is slower in any way that matters. It's four small, boring, individually reversible steps instead of one dramatic one, and every single one of them can survive a rollback cleanly, because none of them ever puts the schema and the currently-running code out of sync with each other. The uncomfortable general lesson underneath the specific one: a rollback plan that's never actually been exercised isn't a rollback plan, it's a guess, and the first time anyone finds out whether it actually works is exactly the worst possible moment for the answer to be no.