Python105 min total · 18 parts
Python Fundamentals for Interviews: Data Structures, Comprehensions, and Gotchas
Part 5 of 18 · ~5 min
Generators and Iterators In Depth
The very first bug flagged in the introduction was open(path).readlines() — reading an entire log file into memory before doing anything with it. A generator fixes it directly:
def read_log_lines(path):
with open(path) as f:
for line in f:
yield line.strip()
Nothing about read_log_lines builds a list anywhere. Each yield hands one line to whatever's asking for it and then stops dead until asked again, which is a completely different shape from a list comprehension's "compute the entire thing, then let you start looking at it." [line for line in open(path)] on a multi-gigabyte log needs the whole file sitting in memory before the first character of output exists; read_log_lines never has more than a single line in hand at once, no matter how big the file behind it is. What you give up for that is replay — there's no going back to line one short of calling read_log_lines again from scratch, and that limitation is about to cause a real bug in a few paragraphs.
How yield actually works
Writing read_log_lines("gateway.log") does not open a file. It builds a generator object and stops — none of the function's own code has run yet, not even the open() call. The body only wakes up, and only runs as far as the next yield, the moment something on the outside asks for a value, whether that's an explicit next() or a for loop pulling its first item:
def read_log_lines(path):
print(f"opening {path}")
with open(path) as f:
for line in f:
print("about to yield a line")
yield line.strip()
lines = read_log_lines("gateway.log") # prints nothing yet — the function hasn't run at all
next(lines) # prints "opening gateway.log", "about to yield a line" — returns line 1
next(lines) # prints "about to yield a line" — returns line 2; file position was remembered
Between calls, everything — the open file handle, the loop's current position — stays frozen exactly where it was. Compare that to an ordinary function, which forgets its entire local state the instant it returns.
Now here's the mistake that catches almost everyone the first time they use a generator for something real. Say the watcher needs two numbers from one pass over a huge log: how many requests errored, and the total time spent across all of them. It's tempting to pull both off the same generator:
lines = read_log_lines("gateway.log")
entries = (to_entry(line) for line in lines) # a generator expression
error_count = sum(1 for e in entries if e.status >= 500)
total_ms = sum(e.duration_ms for e in entries)
print(error_count, total_ms) # 340 0 <- total_ms is ZERO, not wrong, empty
total_ms comes back zero. Not incorrect — empty. error_count's sum() already walked entries from front to back computing as it went, and by the time it returned there was nothing left for total_ms's sum() to consume. There's no rewinding and no hidden second copy — just an exhausted iterator quietly producing nothing at all. A list would have handed you a wrong-but-present number if you accidentally iterated it twice; a generator gives you the right number exactly once and then silence. If you genuinely need two passes, materialize it — entries = list(entries) — or, better, compute everything you need in a single walk.
Generator expressions
Trade a comprehension's square brackets for round ones and the result is a generator instead of a list — that's what both sum() calls above are actually built from. And when a generator expression is the entire argument list of a call, Python lets the function's own parentheses double as the generator's, so sum(e.duration_ms for e in entries) needs only one pair of parens, not the two you'd expect from writing sum((e.duration_ms for e in entries)) in full. Add a second argument next to it and the shortcut goes away — the generator needs its own parentheses again to keep it from being mistaken for two separate arguments.
The iterator protocol, briefly
A for loop has no special knowledge of lists, dicts, strings, or generators individually — it only knows how to work with anything implementing two methods, together called the iterator protocol. __iter__ hands back an iterator (an object able to be stepped through), and repeatedly calling that iterator's __next__ produces the sequence's values one at a time until it raises StopIteration to signal there's nothing left. read_log_lines satisfies that protocol automatically the moment it uses yield — a generator function is essentially a shortcut for writing a class with both of those methods by hand, which is why for line in read_log_lines(path) works at all, and why the exact same loop syntax works identically over a list or a dict without either of them having anything in common with a generator internally.
This is also the honest fix for the "flatten three log files" problem from the comprehensions chapter, once the files are too big to have already been read into lists. itertools.chain takes several iterables — generators included — and walks them one after another as a single iterator, without ever holding more than one line from any of them in memory at once:
from itertools import chain
todays_lines = chain(read_log_lines("morning.log"), read_log_lines("afternoon.log"), read_log_lines("evening.log"))
for line in todays_lines: # still one line in memory at a time, across all three files
...
The nested comprehension from two chapters ago and chain do the same conceptual job — line up several sources and walk them as one — but only chain keeps the lazy, one-line-at-a-time property that makes a generator worth using in the first place.