Skip to main content
CodeOath
← All posts

Python105 min total · 18 parts

Python Fundamentals for Interviews: Data Structures, Comprehensions, and Gotchas

Part 9 of 18 · ~4 min

Mutability, Identity, and is vs ==

These ask two completely different questions. == asks whether two things have the same content. is asks whether two names currently point at one and the same object sitting in memory — not a similar object, the identical one. Confuse the two and the resulting bug has a nasty habit of testing "fine" for a while, right up until the specific values involved change shape:

status_a = int("200")
status_b = int("200")
status_a == status_b    # True — same value
status_a is status_b    # True, on stock CPython — 200 falls inside the small-int cache (-5 to 256)

status_c = int("401")
status_d = int("401")
status_c == status_d    # True
status_c is status_d    # False — 401 is outside that cache

If the watcher only ever checked status is 200, that would look completely reliable for months — every successful request would compare True under is, and nothing would tell you that you were leaning on an implementation detail instead of a language guarantee. The bug shows up the day someone writes if status is 401 to catch the credential-stuffing attempts specifically, and it silently never fires, because 401 was never in the cache range to begin with. == was never the slower option here — it's simply correct at every status code, cached or not, while is happens to work for exactly one narrow range of small integers and gives you no warning when you've stepped outside it.

Reserve is for exactly one job: comparing against None, True, or False. Each of those three is a singleton in CPython — the interpreter builds exactly one of each when it starts up and every part of your program that mentions None, say, is pointing at that same object, never a fresh copy — so identity and equality land on the same answer there every time, with no exceptions to worry about. Everywhere else — status codes, IPs, durations, entries, anything with an actual value worth comparing — == is not a fallback for when is "doesn't apply"; it's simply the correct operator, unconditionally, because it's the only one of the two that was ever promising to compare values in the first place.

Mutable vs. immutable, and why it matters for function arguments

Calling quarantine_all(watchlist) doesn't clone watchlist and hand the clone to the function — it makes ips, inside the function, point at the identical object watchlist already points to outside it. Whether that fact ever surfaces depends entirely on what happens to ips next:

def quarantine_all(ips):
    ips.append("198.51.100.7")     # mutates the SAME list the caller passed in

def replace_list(ips):
    ips = ["0.0.0.0"]              # rebinds the LOCAL name only — caller's list is untouched

watchlist = ["203.0.113.9"]
quarantine_all(watchlist)
watchlist   # ["203.0.113.9", "198.51.100.7"] — the caller's list WAS changed

replace_list(watchlist)
watchlist   # unchanged — reassignment inside the function never propagates back out

Call .append, .sort, or assign through dict[key] = x and you're editing the object itself, in place — so whichever names happen to be pointing at it, anywhere in the program, all observe the same updated object afterward. Reassignment is a completely different operation: it just aims one specific local name at something else, and every other name still pointing at the original object never finds out.

Shallow copy vs. deep copy

Before merging a new batch into the watcher's running by_ip grouping, it's tempting to snapshot the old state first to compare against later:

import copy

by_ip = {"198.51.100.7": [entry_1, entry_2]}
snapshot = by_ip.copy()                # shallow — the inner lists are shared, not duplicated
deep_snapshot = copy.deepcopy(by_ip)

by_ip["198.51.100.7"].append(entry_3)
snapshot["198.51.100.7"]               # now ALSO has entry_3 — same underlying list
deep_snapshot["198.51.100.7"]          # unaffected — every level was recursively copied

.copy() only ever duplicates the one container you called it on — the dict itself gets new storage, but every value sitting inside that storage is copied by reference, so a list living two levels down in by_ip is the same list in both the original and the "copy." copy.deepcopy is the only one of the two that walks all the way down and duplicates every nested container it finds along the way, which is the only way to end up with something genuinely independent. Anyone who's called .copy() expecting full independence and then watched a change to one "copy" bleed into the other has run into precisely this.

That same idea explains a smaller surprise back where this chapter started: a tuple is only immutable one level deep.

entry_with_tags = (LogEntry("198.51.100.7", "POST", "/api/login", 401, 42), ["brute-force", "watching"])
entry_with_tags[1].append("escalated")     # totally legal
entry_with_tags[1]                          # ["brute-force", "watching", "escalated"]
entry_with_tags[0] = None                   # TypeError: 'tuple' object does not support item assignment

What a tuple actually locks is its own slots — you can never make entry_with_tags[1] point at a different list. It says nothing at all about whether the object already sitting in that slot can be mutated internally. A tuple of immutable things (numbers, strings, other tuples) is genuinely frozen all the way through; a tuple holding a list is only frozen at the outer layer, and it's easy to assume "immutable" covers the whole structure when it only ever covered the references.