Docker105 min total · 19 parts
Docker Fundamentals: Images, Containers, and Writing a Good Dockerfile
Part 14 of 19 · ~3 min
Health Checks and Restart Policies
All Docker actually tracks is whether snapledger's main process is still alive — whether it's doing its job is a completely separate question, and not one Docker has any built-in way to answer. A worker stuck retrying the same broken job forever, a connection pool that's quietly exhausted, an API returning nothing but 500s: from where Docker is standing, every one of those is indistinguishable from a healthy container going about its business.
You find this out the specific way that makes it memorable. One customer's receipt — a slightly malformed PDF — sends worker's parsing logic into a loop it never recovers from. The process doesn't crash. docker ps shows it Up, cheerfully, for six hours straight, while every other job in the queue behind that one just sits there, unprocessed, and nothing in Docker's own view of the world says anything is wrong.
An explicit health check is what closes that gap:
HEALTHCHECK --interval=30s --timeout=3s --start-period=10s --retries=3 \
CMD curl -f http://localhost:8080/health || exit 1
web already has a /health endpoint for this. worker doesn't have an HTTP server at all, so its health check has to work differently: every time it successfully finishes a job, it touches a small heartbeat file, and the health check just asks whether that file is suspiciously old.
HEALTHCHECK --interval=30s --timeout=3s --retries=3 \
CMD find /tmp/heartbeat -mmin -2 || exit 1
// worker.js — after each job completes successfully
fs.writeFileSync("/tmp/heartbeat", String(Date.now())); // proves forward progress, not just "still running"
A worker that's stuck reprocessing the same malformed receipt over and over stops touching that file the moment it enters the loop, even though the process itself never exits. docker ps now has a real signal to report — healthy, unhealthy, or starting, in place of a bare Up that never changes — and that's the signal Compose's service_healthy condition is actually built to watch for, the same way a Kubernetes readiness probe or a load balancer's own health check would be. The stuck worker flips to unhealthy within a couple of missed intervals, instead of sitting there quietly for six hours.
Restart policies decide what happens next, once that process has actually stopped:
docker run --restart unless-stopped snapledger-worker
| Policy | What it actually does |
|---|---|
no (default) | Leaves a dead container dead — nothing restarts it automatically |
on-failure[:max-retries] | Restarts it, but only when it exits with an error code, and only up to a limit you can set |
always | Brings it back no matter how it went down, and again after the daemon itself comes back up |
unless-stopped | Same as always, except it honors a docker stop you ran yourself — the daemon won't override that |
unless-stopped is what both web and worker run with in production, and the distinction from plain always matters more than it looks the first time the hosting provider reboots the server for maintenance with no warning. Everything with unless-stopped comes back up on its own, exactly as it was, without anyone touching a keyboard. A third container on the same host — a one-off debug instance you'd deliberately stopped the day before and genuinely didn't want running — would come back too under always, quietly, and you wouldn't find out until you went looking for why something you'd shut off was somehow running again. unless-stopped remembers your last explicit docker stop across the reboot; always does not.