Python105 min total · 18 parts
Python Fundamentals for Interviews: Data Structures, Comprehensions, and Gotchas
Part 13 of 18 · ~2 min
Classes and the Object Model
class Report:
SOURCE = "api-gateway" # class attribute — shared by every Report
def __init__(self, entries):
self.entries = entries # instance attribute — unique per report
morning = Report(morning_entries)
morning.SOURCE # "api-gateway" — found on the class, not the instance
self looks reserved the way def or class is, and it isn't — Python has no idea self means anything special. What's actually happening is that morning.summary() is translated into Report.summary(morning) before it runs, and whichever parameter comes first in summary's definition receives morning automatically, whatever that parameter happens to be named. Rename it to this or me and every method still works exactly the same; the entire language treats it as an ordinary parameter with no special status at all. The only thing making self non-negotiable is that every other Python developer expects it, and code that names it anything else is technically legal and immediately confusing to read.
The class-attribute mutable-default trap
The exact same "evaluated once" bug from two chapters ago comes back one more time, wearing a class's clothes instead of a function's:
class Report:
flagged = [] # DANGER — one list shared by every Report ever created
def flag(self, ip):
self.flagged.append(ip)
morning = Report()
afternoon = Report()
morning.flag("198.51.100.7")
afternoon.flagged # ["198.51.100.7"] — an entry nobody told afternoon about, because it's the identical list object
The cure is the same one from the function-default bug two chapters back, just applied to a class body instead: move the list's creation into __init__ so it's built fresh, per instance, at construction time rather than once, shared, at class-definition time:
class Report:
def __init__(self):
self.flagged = [] # a fresh list per instance
By default, every instance quietly carries a __dict__ — its own private dictionary holding whatever attributes it has, which is what makes self.anything = value always work no matter what the class author anticipated. On a class you're going to create thousands of, like LogEntry if it were a plain class instead of a namedtuple, that per-instance dictionary is real memory overhead, paid again for every single row. __slots__ trades that flexibility away on purpose:
class LogEntry:
__slots__ = ("ip", "method", "path", "status", "duration_ms") # no per-instance __dict__ at all
def __init__(self, ip, method, path, status, duration_ms):
self.ip = ip
self.method = method
self.path = path
self.status = status
self.duration_ms = duration_ms
e = LogEntry("198.51.100.7", "POST", "/api/login", 401, 42)
e.geo = "RO" # AttributeError — 'geo' isn't in __slots__, and there's no __dict__ to fall back on
Declaring __slots__ reserves fixed storage for exactly those attribute names and removes the per-instance dictionary entirely, which measurably shrinks memory on a class instantiated a hundred thousand times over a busy log. The trade is that the class becomes genuinely closed — no attribute outside the declared set, ever, on any instance — which is a constraint worth choosing on purpose rather than a default worth reaching for everywhere.