Skip to main content
CodeOath
← All posts

Docker105 min total · 19 parts

Docker Fundamentals: Images, Containers, and Writing a Good Dockerfile

Part 7 of 19 · ~3 min

CMD vs. ENTRYPOINT: The Instruction Everyone Mixes Up

A week after worker ships, you notice every deploy takes an extra ten seconds longer than it should. Digging in, the pattern is specific: docker compose up -d --no-deps worker to roll out a new version always takes exactly ten seconds to stop the old container, no matter how quickly worker.js itself finishes whatever job it was mid-processing.

The Dockerfile has this line:

CMD node worker.js

No brackets — that's the shell form, and it's the whole bug. Docker doesn't run node worker.js directly when it's written this way; it hands the whole string to /bin/sh -c, so the shell is the process the container actually considers its main one, and Node is just something the shell happens to have launched underneath itself. docker stop's SIGTERM has exactly one address to go to — the container's main process — and that address belongs to the shell, not to Node. Whether the shell bothers relaying it to the thing it launched is not guaranteed at all, and here it doesn't, so nothing ever shuts down on its own. Docker sits there for the entire ten-second grace period and then ends both processes with SIGKILL anyway.

The exec form — a JSON array — fixes it, because it skips the shell entirely and runs your process directly as PID 1, where it receives signals like SIGTERM itself:

CMD ["node", "worker.js"]
docker run snapledger-worker                 # runs: node worker.js, as PID 1, receiving signals directly

Once worker.js is actually getting the signal, it can listen for it and shut down on its own terms — finish the job it's mid-way through, close its queue connection, then exit — instead of being killed off mid-write:

process.on("SIGTERM", async () => {
  await finishCurrentJob();   // let whatever receipt is mid-parse actually complete
  await queue.disconnect();   // close the connection cleanly, don't just drop it
  process.exit(0);
});

time docker stop worker goes from real 0m10.043s — the full grace period, every single time, force-killed at the end of it — to real 0m0.312s, because the container now exits on its own the moment it's actually done, instead of Docker having to wait out a timeout for a process that was never going to respond.

ENTRYPOINT is the other half of this, and the two only really click once you see them paired. Whatever CMD names by itself is just a suggestion, entirely at the mercy of whoever types docker run — pass anything after the image name and CMD is gone, no trace. ENTRYPOINT refuses to be displaced like that: it's the command the container runs no matter what, and anything you tack onto docker run doesn't replace it, it gets tacked onto the end of it instead. A month later, when you need to run a one-off backfill — reparse every receipt from the last thirty days without starting the normal queue loop — this pairing is exactly what makes it clean:

ENTRYPOINT ["node", "worker.js"]
CMD ["--loop"]
docker run snapledger-worker                 # runs: node worker.js --loop   (the normal queue worker)
docker run snapledger-worker --backfill       # runs: node worker.js --backfill   (only CMD's part changes)

ENTRYPOINT answers "what is this image, fundamentally," and CMD answers "and what should it do if nobody says otherwise" — two separate questions, each with its own owner. Whoever calls docker run only ever gets a say over the second one, which means they can steer the container into a different mode without ever having to know, retype, or risk mistyping the actual command underneath. That split is the shape you reach for anytime an image should behave like a single fixed tool with a switchable mode — which, once you notice it, is most of what actually gets built.