Docker105 min total · 19 parts
Docker Fundamentals: Images, Containers, and Writing a Good Dockerfile
Part 2 of 19 · ~3 min
What Problem Docker Actually Solves
You've been shipping snapledger — a receipt-scanning expense tracker for small businesses — off three laptops for four months, and it's mostly worked because the three of you keep talking to each other. That stops scaling the moment a fourth machine enters the picture that nobody is personally maintaining: a staging server, and soon after it, a customer's production server. Nobody is going to SSH in every week and manually keep its Node version, its system libraries, and its environment variables in sync with whatever the three of you happen to be running that day. That kind of manual alignment is exactly what breaks — quietly, and always at the worst time.
Docker's answer is to stop treating "the code" and "everything the code needs to run" as separate concerns. Both get bundled into one shippable thing — an image — and running that same image is what happens, indistinguishably, on your laptop, Priya's laptop, a CI runner, or a customer's server you'll never personally log into. The host machine's own OS, its installed packages, whatever Node version happens to be on its PATH — none of it matters anymore, because the thing you're running carries its own copy of all of that with it.
It's tempting to file this next to virtual machines, since they sound like they're solving the same problem, but the two approaches are structurally different:
| Virtual Machine | Container | |
|---|---|---|
| What's actually virtualized | The whole machine — its own kernel included | One process, running on the kernel the host already has |
| Time to get one running | Whole seconds at best, often a minute or more — it's booting an OS | Under a second, most of the time |
| Typical footprint on disk | Gigabytes, because a full OS ships with it | Usually well under a gigabyte |
| How solid the isolation is | About as solid as it gets — a genuinely separate kernel | Solid enough for most purposes, but resting on kernel features rather than a hard wall |
| How many fit on one host | Not many — each one drags its own OS along for the ride | A lot — they're splitting one kernel between them |
The distinction worth actually holding onto: a container is not a stripped-down virtual machine. Strip away the marketing and it's an unremarkable process, running on the same kernel as everything else on the box, with two Linux kernel features doing all the work of convincing it — and you — that it's alone there. One of those, namespaces, is what a container's isolation is actually built from: swap out what a process can see for the process list, the network interfaces, the mounted filesystems, and the hostname, and there's simply nothing outside that view for it to bump into. The other, cgroups, is a metering and enforcement layer sitting on top — it decides how much CPU, memory, and disk/network I/O that process gets to touch, and keeps counting as it runs. Skip booting a second kernel entirely, and startup drops from the better part of a minute to a handful of milliseconds, with almost none of a VM's overhead along the way. The cost shows up on the security side of the ledger instead: escape a namespace and there's a real kernel directly on the other side of it, not another layer of isolation, which is a boundary a fully separate VM kernel is generally far better positioned to hold.
That startup-time gap isn't an academic footnote for snapledger. Months from now, when Berrywell's statements are backing up and you need three more worker instances right now to clear the queue, the difference between "under a second" and "the better part of a minute, per instance, if you'd built this on VMs instead" is the entire difference between the backlog clearing before anyone notices and a customer opening a support ticket while you wait for machines to boot.