Python105 min total · 18 parts
Python Fundamentals for Interviews: Data Structures, Comprehensions, and Gotchas
Part 10 of 18 · ~5 min
Decorators
Parsing a very large log file is the watcher's slowest step, and knowing exactly how slow is worth a permanent, reusable answer rather than a one-off time.perf_counter() sprinkled through the code:
import functools
import time
def timer(func):
@functools.wraps(func) # preserves func's __name__, docstring, etc. on the wrapper
def wrapper(*args, **kwargs):
start = time.perf_counter()
result = func(*args, **kwargs)
elapsed = time.perf_counter() - start
print(f"{func.__name__} took {elapsed:.4f}s")
return result
return wrapper
@timer
def parse_batch(path):
return [to_entry(line) for line in read_log_lines(path)]
parse_batch("gateway.log")
# "parse_batch took 0.0421s"
Read the @timer line as nothing more than parse_batch = timer(parse_batch), run immediately after the function is defined — the name parse_batch ends up bound to whatever timer handed back, which is wrapper, not the original function. wrapper can accept any function's arguments specifically because it was written with *args, **kwargs rather than a fixed signature, the exact trick from the previous chapter.
Leaving out @functools.wraps(func) doesn't break anything the watcher actually runs — parsing still works — but it swaps the wrapped function's identity out from under it. parse_batch.__name__ becomes "wrapper", parse_batch.__doc__ becomes whatever docstring (if any) wrapper has, and every tool downstream that reports on a function by inspecting these — a logger, a profiler, an API framework generating docs from function names — starts reporting on the wrong function without raising so much as a warning.
Decorators with arguments
The watcher is eventually going to call a slow, occasionally-flaky geolocation API (that's the whole last chapter), and a flaky network call needs a retry, not just a timer:
def retry(times):
def decorator(func):
@functools.wraps(func)
def wrapper(*args, **kwargs):
last_error = None
for attempt in range(times):
try:
return func(*args, **kwargs)
except Exception as e:
last_error = e
print(f"Attempt {attempt + 1} failed: {e}")
raise last_error # re-raise the LAST real failure — see below
return wrapper
return decorator
@retry(times=3)
def lookup_geo(ip):
... # built out fully in the GIL/async chapter
There's an extra step hiding in @retry(times=3) that @timer never needed. Python evaluates retry(times=3) first — an ordinary function call, nothing to do with decoration yet — and only the result of that call, decorator, ends up applied to lookup_geo. retry itself never touches lookup_geo at all; it exists purely to build and hand back the function that will.
Leave the parentheses off — @retry instead of @retry(times=3) — and this fails, but not where you'd expect. retry runs immediately with lookup_geo standing in for its times parameter, and dutifully returns decorator, which then gets bound to the name lookup_geo. No error yet — decoration "succeeds." Call lookup_geo(some_ip) and you're really calling decorator(some_ip), which returns yet another wrapper instead of a result — still no error. Only when that gets called does for attempt in range(times) finally run with times still holding the original function object instead of a number, and TypeError: 'function' object cannot be interpreted as an integer shows up two calls away from the line that actually caused it.
Notice the loop raises last_error explicitly rather than a bare raise after it finishes. That distinction matters more than it looks: a bare raise only works inside an except block, where it re-raises whatever's currently being handled. Once the for loop runs out of attempts and falls through normally, Python is no longer inside any except block — the last one already exited — so a bare raise there fails with an unrelated error of its own, RuntimeError: no active exception to reraise, instead of the ConnectionError or whatever actually went wrong three attempts ago. Track the real failure in a variable and raise that.
Built-in decorators worth knowing: @property, @staticmethod, @classmethod
class Report:
def __init__(self, entries):
self.entries = entries
@property
def error_rate(self):
if not self.entries:
return 0.0
errors = sum(1 for e in self.entries if e.status >= 500)
return errors / len(self.entries) # looks like an attribute, costs a method call
@staticmethod
def parse_line(line):
_, ip, method, path, status, duration = line.split()
return LogEntry(ip, method, path, int(status), int(duration)) # doesn't need self or cls
@classmethod
def from_file(cls, path):
entries = [cls.parse_line(line) for line in read_log_lines(path)]
return cls(entries) # cls, so a subclass gets back ITS OWN type
report = Report.from_file("gateway.log")
report.error_rate # 0.0183 — no parentheses, reads like a plain attribute
Three different jobs, easy to blur together. error_rate genuinely runs code every time it's read — dividing the current error count by the current total — but nothing at the call site (report.error_rate, no parentheses) betrays that; @property is the mechanism that lets a computed value masquerade as a stored one, so the class is free to change how error_rate is derived later without every caller needing to change report.error_rate() into report.error_rate. parse_line never touches self, because it doesn't need an instance at all to turn one raw string into a LogEntry — it's to_entry from the comprehensions chapter, relocated onto the class it's most related to, and @staticmethod is what tells Python not to bother passing it an instance it would never use. from_file is different again: it needs to produce a Report, but written against cls rather than the literal name Report, so a subclass calling ErrorReport.from_file(path) gets back an ErrorReport, not a plain Report pretending to be one.
One built-in decorator is worth knowing before writing another custom one: functools.lru_cache is the @timer-shaped decorator's caching cousin, and the async chapter is about to call the same geolocation lookup for the same handful of repeat offenders more than once:
@functools.lru_cache(maxsize=None)
def lookup_geo_cached(ip):
... # the real, slow lookup — runs at most once per distinct ip
lookup_geo_cached("198.51.100.7") # runs the real lookup
lookup_geo_cached("198.51.100.7") # returns the cached result instantly, doesn't run the body again
It only works because its argument is hashable — the same property that let a tuple, but never a list, sit inside a set back in the first chapter — which is what lru_cache uses internally to recognize "this call again."