Skip to main content
CodeOath
← All posts

Docker105 min total · 19 parts

Docker Fundamentals: Images, Containers, and Writing a Good Dockerfile

Part 19 of 19 · ~4 min

Common Mistakes Worth Remembering

Every one of these actually happened somewhere in the story above.

  • COPY . . before installing dependencies — this is what made Priya's one-line CSS fix trigger a full npm ci every time, before the Dockerfile's copy order got fixed.
  • Cleaning up a big file in a RUN that comes after the one that wrote it — the fixture archive's bytes had already been sealed into an earlier layer; only doing the cleanup inside that same RUN ever actually shrinks anything.
  • Writing CMD/ENTRYPOINT in the shell-string form for something that needs to shut down cleanly — this is exactly why every worker deploy was taking ten extra seconds: the shell ended up as the container's real process, and it never bothered passing SIGTERM down to the thing it had launched.
  • Trusting depends_on by itself to mean Postgres is ready for queries, when all it actually promises is that the container was launched — pair it with a healthcheck and condition: service_healthy, or you get the three-restarts-on-first-run bug.
  • Baking an environment-specific value into an image with ENV instead of injecting it at runtime — this is exactly how Priya's staging-only DATABASE_URL ended up tagged 1.3.0 indistinguishably from every other production build, and how production traffic pointed at the staging database for eleven minutes.
  • Copying node_modules across build stages with different base images — the Berrywell incident happened because the multi-stage build's compile stage used node:20 (glibc) while the runtime stage used node:20-alpine (musl), and a compiled dependency's native binary silently didn't match.
  • Letting a secret ride into a layer, whether that's through ARG, ENV, or a stray .env file COPY . . swept up without anyone meaning it to — this nearly happened twice, once with a Stripe key as a build argument and once with Marcus's local .env.
  • Assuming anything written inside a container will survive that container — nothing does, unless it's sitting in a volume. pgdata living as a named volume rather than as part of the container is the entire reason recreating Postgres didn't erase every expense record snapledger had.
  • Assuming a smaller base image is a drop-in, risk-free swap — Alpine's size advantage is real, but it comes from a different C library, and the first input dense enough to exercise a compiled dependency's native path is exactly when that difference turns into an outage.
  • Leaving resource limits off, or assuming a memory limit and a CPU limit fail the same way — they don't. A memory limit kills suddenly and loudly (exit 137); a CPU limit throttles silently, and the two incidents in this story looked completely different because of it.
  • Reaching for a bigger base image than the job needs, or forgetting .dockerignore entirely — the fixtures folder, the full .git history, and very nearly a .env file all rode along into the build context before anyone thought to exclude them.
  • Leaving every version unpinned. FROM node:latest is a moving target, not a fixed one — the image behind that tag can change out from under you at any point, and when it does, the whole premise of "this runs the same everywhere" quietly stops being true, without a single line of your own code changing to explain why.

None of these were exotic mistakes. Every one of them is the kind of thing that looks completely reasonable in the moment — a Dockerfile written in the order the instructions happened to occur to someone, a base image swapped for a smaller one because smaller is obviously better, a config value hardcoded because it was Friday and the "proper" way would take twenty more minutes. Docker doesn't stop you from doing any of it. It just makes the consequences specific and traceable instead of vague — a layer you can point to, an exit code you can look up, a .State field with the actual answer sitting in it — which is most of what separates "we have no idea why this broke" from a five-minute fix.

Everything above stops at the container boundary — what actually moves snapledger-web's image from a laptop to a customer's server, on a schedule, without someone typing docker push by hand every time, is a separate machine entirely, and CI/CD Pipelines Explained is where that machine gets taken apart. Once web, worker, redis, and postgres stop being four containers in one small team's Compose file and start being four things different teams own, the split between them turns into an organizational question as much as a technical one — Microservices vs. Monolith is where that question gets its own treatment. For something closer to a real, working Compose setup than anything a receipt-scanning example can show on its own, this site's infra/ folder runs an actual self-hosted code-execution engine behind one. And if reading about docker build and docker exec isn't the same as having typed them, the code lab is where that gap actually closes.