CI/CD & DevOps65 min total · 17 parts
CI/CD Pipelines Explained: From Push to Production
Part 13 of 17 · ~2 min
Deployment Strategies
worker's deploys start out as the simplest possible thing — stop the old container, start the new one — and the ten-second SIGTERM-to-SIGKILL gap from the Docker chapters is exactly why that's noticeable: for a few seconds, nothing is pulling jobs off the queue at all, and on a slow evening nobody minds.
| Strategy | The mechanism | Getting back out | What it costs |
|---|---|---|---|
| Recreate | Kill the current container, bring up the replacement | Slow — the old build has to be stood back up from nothing | Nothing extra, and it's what the team starts with |
| Rolling | Swap instances in small batches while the rest keep serving | Moderate — halt partway through and reverse direction | Nothing extra |
| Blue-green | Stand up the new version fully, then flip traffic to it all at once | Instant — flip traffic back | The whole fleet, doubled, for as long as the old one sticks around |
| Canary | Let a thin sliver of real users hit the new build first, widen it gradually | Fast — pull the sliver back out of rotation | Real infrastructure and monitoring to split traffic by version at all |
Once Berrywell's backlog forces the team to run three worker instances instead of one — the exact scaling story from the Docker chapters — recreate stops making sense on its own, because stopping all three at once means the whole queue goes silent, not just briefly slower. Rolling updates become the natural next step: replace one instance, confirm it's healthy, move to the next, and the queue never fully stops. The trade nobody thinks about until it bites them is what happens to the database in between: for however long the rollout takes, both the old and new worker code are querying the same tables simultaneously, so any schema change made during that window has to satisfy both versions at once, not just the one about to become current — a constraint that turns out to matter a great deal one chapter from now.
web, being what a customer actually clicks through and types payment-adjacent data into, eventually graduates to canary: a handful of real sessions land on a new build first, for a few minutes, before it's trusted with everyone. It costs more to run — someone has to actually split traffic and watch metrics per version — but it's the only one of the four willing to accept a little live exposure on purpose, specifically so a bad build gets caught by a handful of unlucky requests instead of every single one at once.