Python105 min total · 18 parts
Python Fundamentals for Interviews: Data Structures, Comprehensions, and Gotchas
Part 4 of 18 · ~3 min
Comprehensions In Depth
Turning a batch of raw lines into LogEntry objects is the watcher's most common operation, and it's where comprehensions earn their keep first:
def to_entry(line):
_, ip, method, path, status, duration = line.split()
return LogEntry(ip, method, path, int(status), int(duration))
entries = [to_entry(line) for line in raw_lines] # list comprehension
slow_requests = [e for e in entries if e.duration_ms > 200] # filter
status_of = {e.ip: e.status for e in entries} # dict comprehension
unique_paths = {e.path for e in entries} # set comprehension
The speed difference isn't imaginary, either: a for loop calling .append() pays for a full method lookup and a Python-level function call on every single pass, while a comprehension's looping is handled by the interpreter's own bytecode with none of that overhead — measurably quicker on anything but a tiny list. That advantage stops paying for itself the moment a comprehension needs a second for clause or more than one condition crammed onto one line; past that point the equivalent loop is easier to step through in a debugger, and being easy to step through is worth more than the saved microseconds.
Nested comprehensions, and the order that trips people up
A real gateway rotates its log into a new file every few hours, so "today's traffic" is usually several files, not one:
todays_batches = [morning_lines, afternoon_lines, evening_lines]
all_lines = [line for batch in todays_batches for line in batch] # flatten three files into one list
Whichever for comes first in the comprehension is the loop that would be written first if you unrolled it by hand — for batch in todays_batches is on the outside, for line in batch is nested underneath it, exactly matching the reading order left to right. Reverse them, writing for line in batch for batch in todays_batches, and Python doesn't just get confused about what you meant — it hits NameError: name 'batch' is not defined, because at the point the outer-looking clause runs, nothing has introduced batch yet. That specific error, rather than a silently wrong answer, is usually what tips people off that the order is backwards.
This flattening trick works fine for three lists already sitting in memory. It stops being the right tool the moment those "batches" are themselves generators — three separate calls to read_log_lines, say — because a nested comprehension over them would need to hold all three open at once. The generators chapter has the better answer for that case.
The conditional expression vs. the filter clause
Put the word "if" in two different positions relative to "for" and you get two operations that share almost no behavior:
# Filter clause — keep only the failed logins, DROP everything else
failed_logins = [e for e in entries if e.path == "/api/login" and e.status == 401]
# Conditional expression — label EVERY entry, drop nothing
labeled = ["FLAG" if e.status >= 500 else "ok" for e in entries]
failed_logins can come out shorter than entries — every item that fails the if never makes it into the result at all. labeled is guaranteed to come out exactly as long as entries, always, because the if/else sitting before for isn't deciding whether to keep an item, it's deciding what value to compute for it. One line can shrink the collection; the other line can't, no matter what condition you write into it.