Docker105 min total · 19 parts
Docker Fundamentals: Images, Containers, and Writing a Good Dockerfile
Part 10 of 19 · ~4 min
Choosing a Base Image
With staging looking solid, the team starts thinking about what actually goes into production. FROM is one line, and it ripples further than any other single choice in a Dockerfile:
| Base type | Example | Roughly how big | What you're trading away |
|---|---|---|---|
| Full distribution | node:20, ubuntu:22.04 | Hundreds of MB and up | Nothing — full shell, package manager, every debugging tool you'd expect |
| Slim | node:20-slim | A fraction of the full size | Still Debian and still glibc, just with most non-essential packages stripped out |
| Alpine | node:20-alpine | A fraction of slim again | A different C library (musl) under the hood — occasionally incompatible with something compiled against glibc |
| Distroless | gcr.io/distroless/nodejs20 | As small as it practically gets | Strips out the shell entirely, so there's simply nothing for docker exec to open |
Alpine isn't just Debian with fewer things installed — underneath it is swapping out the C standard library itself, musl in place of glibc, which is a far more fundamental change than trimming a package list. And it's exactly that swap that turns into a real production bug for you, two weeks after launch.
A customer — a small grocery chain called Berrywell — starts uploading monthly statement PDFs that run to fifty or sixty pages, much larger than anything the team tested with locally. The very first one Berrywell uploads takes down the worker's parsing job with a cryptic native-module error, on an image that had worked flawlessly for every smaller receipt from every other customer for two weeks straight.
The root cause is exactly the plant from the multi-stage chapter: worker's build stage is node:20 (glibc), and one of its dependencies — the PDF-rendering library used to split a multi-page statement into individual page images — ships a compiled native binary. For small, simple receipts that library happened to fall back to a pure-JavaScript path that never touched the compiled binary at all. A dense, sixty-page scanned statement is the first input that actually exercises the fast native path — and that binary was linked against glibc in the build stage, then copied straight into node:20-alpine's musl-based runtime stage.
The failure is stranger than a normal crash, and it's worth knowing the shape of it so it's recognizable next time: the file is right there, ls sees it, the permissions are fine, and trying to run it anyway reports back something like no such file or directory — which is a flatly wrong description of what's actually happening. What's missing isn't the binary. It's the dynamic linker the binary was compiled to ask for, at a path like /lib64/ld-linux-x86-64.so.2, which simply doesn't exist on a musl-based image. The kernel goes looking for the interpreter a glibc binary declares it needs, doesn't find it anywhere on the filesystem, and reports the whole thing as "not found" — not "wrong linker," which is the actual problem and the thing that makes this specific error so easy to misdiagnose as a missing file rather than a mismatched one.
The fix is to make sure the environment a native dependency compiles in matches the environment it runs in — build on Alpine if you're going to run on Alpine:
FROM node:20-alpine AS build
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build
FROM node:20-alpine
WORKDIR /app
COPY --from=build /app/dist ./dist
COPY --from=build /app/node_modules ./node_modules
CMD ["node", "dist/worker.js"]
The lesson generalizes past this one library. Going smaller is still worth doing — fewer packages sitting in the image means fewer packages that could ever turn up in a CVE scan, and every worker instance you spin up to burn down a backlog gets to serving traffic that much sooner because there's simply less to download before it can start. None of that is in question. What Berrywell's statement proved is that "smaller" and "identical" are two different promises, and the gap between them tends to hide exactly where nobody's been testing — the corner case big enough, or unusual enough, that it's the first thing to actually exercise the part of the app that changed.
Distroless pushes the same trade even further, dropping the shell and the package manager entirely for a footprint smaller than Alpine's. You seriously consider it for worker once Berrywell's incident is behind you and the image is stable — and decide against it, for a reason that has nothing to do with size: no shell means no docker exec -it worker sh, and a production worker with an intermittent bug is exactly the kind of thing you expect to need to get inside of, live, at an inconvenient hour. That trade-off is worth naming explicitly rather than defaulting into: distroless is the right call for a runtime nobody ever needs to debug interactively, and the wrong one the moment "getting a shell into this container at 2 a.m." is a realistic thing you'll need to do — which, a few chapters from now, it is.