Docker105 min total · 19 parts
Docker Fundamentals: Images, Containers, and Writing a Good Dockerfile
Part 18 of 19 · ~3 min
Resource Limits
Skip memory and CPU limits entirely, and there's nothing stopping one runaway container from claiming every last bit of CPU and RAM the host has — with everything else on that machine, containerized or not, left to fight over whatever's left. worker had a limit set from day one, which is exactly why it stayed contained: it went down alone, forty seconds at a time, instead of taking web and Postgres down with it in the same crash.
docker run --memory=512m --cpus=1.5 snapledger-worker
services:
worker:
image: snapledger-worker
deploy:
resources:
limits:
memory: 512M
cpus: "1.5"
The root cause of the 2 a.m. page, once you trace it all the way back, is Berrywell again: their statement PDFs run fifty-plus pages, and worker was rendering every page of a document into memory at once before starting to process any of them. A ten-page receipt never got close to 512MB. A sixty-page one blew straight through it.
Here's the asymmetry that explains why this incident looked nothing like a slowdown you'd chased down the week before. A memory limit is a hard cap — cross it, and the kernel's OOM killer ends the process outright, the same 137 from the last chapter, no warning. A CPU limit works on an entirely different principle: go over it and nothing dies, the scheduler just starts rationing — the container keeps running, only with less CPU time handed to it than it's actually asking for, which reads as things getting gradually slower, never as a crash.
The week before, web had a rough morning during a burst of signups nobody had planned for, and the symptom looked nothing like this incident: no restarts, no exit codes, no alerts — just every request taking noticeably, almost imperceptibly longer to answer than it had the day before, for about ninety minutes, and then it quietly went away on its own once the burst passed. You'd chalked it up to Postgres being busy at the time. Looking back at it now, with docker stats open during a deliberate load test, the pattern is unmistakable: CPU % pinned at a flat, suspiciously exact ceiling the whole time, which is precisely what a CPU limit being hit looks like from the outside — the container asking for more CPU than its --cpus allowance and simply not getting it, over and over, with nothing in any log to say so. Memory failures are loud and sudden; CPU pressure is invisible until you already know to look for it, which is exactly why it took a second, much louder incident to teach you to recognize the quiet one in hindsight.
The actual fix isn't raising the memory limit and hoping — it's changing what worker does:
// before: subgoal "render every page" happened all at once, so peak memory
// scaled with how many pages the receipt had — fine for 10, fatal for 60
const pages = await renderAllPages(pdfBuffer); // every page resident in memory at once
for (const page of pages) await scoreAndStore(page);
// after: one page in memory at a time — peak memory is now roughly constant,
// however many pages the document turns out to have
for await (const page of renderPagesOneAtATime(pdfBuffer)) {
await scoreAndStore(page); // page goes out of scope before the next one renders
}
The limit stays at 512MB. It's the job that changes — and the fix is specifically about peak memory, not total work done, which is why it scales to Berrywell's next statement being ninety pages instead of sixty without anyone having to revisit this chapter again.