Skip to main content
CodeOath
← All posts

CI/CD & DevOps65 min total · 17 parts

CI/CD Pipelines Explained: From Push to Production

Part 13 of 17 · ~2 min

Deployment Strategies

worker's deploys start out as the simplest possible thing — stop the old container, start the new one — and the ten-second SIGTERM-to-SIGKILL gap from the Docker chapters is exactly why that's noticeable: for a few seconds, nothing is pulling jobs off the queue at all, and on a slow evening nobody minds.

StrategyThe mechanismGetting back outWhat it costs
RecreateKill the current container, bring up the replacementSlow — the old build has to be stood back up from nothingNothing extra, and it's what the team starts with
RollingSwap instances in small batches while the rest keep servingModerate — halt partway through and reverse directionNothing extra
Blue-greenStand up the new version fully, then flip traffic to it all at onceInstant — flip traffic backThe whole fleet, doubled, for as long as the old one sticks around
CanaryLet a thin sliver of real users hit the new build first, widen it graduallyFast — pull the sliver back out of rotationReal infrastructure and monitoring to split traffic by version at all

Once Berrywell's backlog forces the team to run three worker instances instead of one — the exact scaling story from the Docker chapters — recreate stops making sense on its own, because stopping all three at once means the whole queue goes silent, not just briefly slower. Rolling updates become the natural next step: replace one instance, confirm it's healthy, move to the next, and the queue never fully stops. The trade nobody thinks about until it bites them is what happens to the database in between: for however long the rollout takes, both the old and new worker code are querying the same tables simultaneously, so any schema change made during that window has to satisfy both versions at once, not just the one about to become current — a constraint that turns out to matter a great deal one chapter from now.

web, being what a customer actually clicks through and types payment-adjacent data into, eventually graduates to canary: a handful of real sessions land on a new build first, for a few minutes, before it's trusted with everyone. It costs more to run — someone has to actually split traffic and watch metrics per version — but it's the only one of the four willing to accept a little live exposure on purpose, specifically so a bad build gets caught by a handful of unlucky requests instead of every single one at once.