Skip to main content
CodeOath
← All posts

Architecture & Patterns110 min total · 26 parts

ACID, SOLID, and Design Patterns: A Complete Software Design Reference

Part 5 of 26 · ~2 min

Durability: Write-Ahead Logs and Crash Recovery

Durability covers what happens after a commit succeeds — specifically, that nothing which happens afterward, including total power loss the instant the acknowledgment reaches the caller, is allowed to undo it. Picture the worst possible timing for it to matter: the payment gateway confirms the charge, the database reports the booking as committed, a confirmation email gets queued up — and the hosting box loses power half a second later, in the middle of the checkout rush. Durability is why that booking is still sitting there, intact, once the box finishes rebooting.

Nearly every engine solves this with a write-ahead log (WAL) — a running, append-only record of intended changes that sits apart from the actual table files. A change gets written to that log first, and the commit isn't reported back to the caller as successful until the log entry itself is confirmed down on durable storage. The real table files can lag behind and catch up whenever it's convenient, because if the process dies before they do, startup recovery works backward from the most recent checkpoint and re-applies whatever the log recorded past that point, one entry at a time, until the data is caught back up to where the commits said it should be.

Without a log: every commit means an in-place write scattered somewhere
inside the real data file. That write is slow, and if the process dies
mid-write, the file is left in an unknown, possibly corrupted state — there's
no record of what was half-written.

With a WAL: every commit means one append to the end of a log file, which
is fast because it never has to seek. The real data file catches up later,
and a crash before that just means replaying entries the log already has.

This is also precisely where durability trades against raw throughput. Most engines expose a setting for how aggressively they flush that log to disk: fsync on every single commit is fully durable and measurably slower under load; fsync on an interval is faster and carries a small window where a handful of recent commits could be lost in a true hard crash. It's a real, explicit knob — not every write genuinely needs the strongest guarantee available, and a box office system deciding "we can tolerate losing the last twenty milliseconds of holds in a true crash, in exchange for handling the on-sale spike" is a defensible, common choice.