Node.js95 min total · 14 parts
Node.js Fundamentals: The Runtime, the Event Loop, and Building Real APIs
Part 8 of 14 · ~2 min
Working with the File System & Streams
The problem with loading the whole receipt into memory
const fs = require("fs");
fs.readFile("./receipts/ord_4471.pdf", (err, data) => {
// `data` is the ENTIRE file, held in memory as one Buffer, all at once
forwardToStorage(data);
});
fs.readFile — and the Promise-flavored version alongside it — won't hand control back until it's pulled the file's full contents into memory; the callback simply doesn't fire before that's done. For a 40KB receipt, that's nothing. For the rare 200MB bulk-export PDF a corporate customer's accounting system attaches, billhook is allocating that much process memory just to relay it onward, and doing this for a handful of large receipts concurrently — which is exactly what happens right after Meridian recovers from an outage and replays a backlog — can exhaust available memory outright.
Streams as the fix
A stream flips that around: it hands you each chunk of data the moment that chunk is available, rather than making you wait for everything to finish first.
const fs = require("fs");
function forwardReceiptToStorage(orderId, res) {
const readStream = fs.createReadStream(`./receipts/${orderId}.pdf`); // chunks, not all at once
readStream.pipe(uploadStreamFor(orderId)); // forwards each chunk as it's read
}
.pipe() is the plumbing between them — a readable source (the file sitting on disk) and a writable destination (the outbound upload) — relaying each chunk across the instant the readable side produces it. Memory use stays roughly constant no matter how large the receipt is — at any given moment billhook is holding exactly one chunk (tens of kilobytes, typically) rather than the file in its entirety.
Backpressure
There's more to piping than shuttling bytes across as quickly as possible, though — it has to solve backpressure: the mismatch that shows up whenever the source can produce faster than the destination can keep up, which for billhook is exactly what happens when a receipt streams off a fast local disk into a comparatively slow upload connection. The signal a writable stream gives for "I'm full, ease off" is its .write() call returning false once its internal buffer crosses some threshold; it announces "I'm ready again" by firing a 'drain' event once that buffer has room.
readStream.on("data", (chunk) => {
const canContinue = uploadStream.write(chunk);
if (!canContinue) {
readStream.pause(); // stop reading until the upload catches up
uploadStream.once("drain", () => readStream.resume());
}
});
All of that pause/resume/drain choreography is precisely what .pipe() handles for you without being asked — which is why billhook reaches for it (or stream.pipeline(), which adds proper error and cleanup handling on top) instead of wiring 'data' and 'write' listeners together by hand. Leave backpressure unhandled in a pipeline you built yourself, and the moment the writable side can't keep pace, memory usage climbs without a ceiling until something gives.