Testing64 min total · 12 parts
Testing Fundamentals: Unit, Integration, and E2E Tests Done Right
Part 9 of 12 · ~4 min
Mocking in Python
The grader runs into the same category of problem the Node service did: dependencies that are slow, external, or both. Python answers this with unittest.mock, usable straight from pytest without pulling in anything extra, centered on two classes — Mock and MagicMock — plus a patch tool for temporarily standing a fake in for a real object during a test.
from unittest.mock import Mock
storage = Mock()
storage.fetch.return_value = Submission(id="sub_91", exercise_id="ex_rate_limiter", code="...")
submission = storage.fetch("sub_91")
assert submission.exercise_id == "ex_rate_limiter"
storage.fetch.assert_called_once_with("sub_91")
Touch any attribute or method on a Mock and it springs into existence on the spot, itself another Mock by default, and every call against it gets logged so you can check later with assert_called_once_with, call_count, or call_args.
MagicMock does everything Mock does, plus it arrives with Python's dunder methods already wired up — __len__, __iter__, __enter__, __exit__, and the rest — and our own grader is exactly why that difference is worth knowing rather than a footnote to skip past. Look at the line in grade that does the actual work: with sandbox_for(submission) as box:.
sandbox_for is used as a context manager, which means Python calls __enter__ on whatever it returns. A plain Mock does not implement __enter__, so substituting one produces TypeError: 'Mock' object does not support the context manager protocol — a confusing failure, because nothing in the traceback mentions the word "mock." MagicMock handles it:
from unittest.mock import MagicMock
fake_sandbox = MagicMock()
# `with sandbox_for(...) as box` binds box to the __enter__ return value, so configure that
fake_sandbox.__enter__.return_value.run.return_value = (0, "")
Same rule for anything used with len(), iterated in a for loop, indexed with [], or unpacked. When a mock fails in a way that seems to have nothing to do with mocking, a missing dunder is the first thing to check.
monkeypatch
pytest ships a monkeypatch fixture that temporarily overwrites an attribute, an environment variable, or a dictionary entry for exactly the lifetime of one test, then puts everything back the way it was on its own — no manual save-then-restore dance required from you. The grader pulls its per-case time limit straight out of the environment, which makes it a natural thing to pin down in a test:
# grader/config.py
import os
def case_timeout_seconds():
return int(os.environ["GRADER_CASE_TIMEOUT_SECONDS"])
def test_case_timeout_comes_from_the_environment(monkeypatch):
monkeypatch.setenv("GRADER_CASE_TIMEOUT_SECONDS", "5")
assert case_timeout_seconds() == 5
# GRADER_CASE_TIMEOUT_SECONDS is restored to whatever it was, once this test ends
def test_fetch_submission_never_touches_the_real_network(monkeypatch):
def fake_get(url, timeout=None):
return Mock(json=lambda: {"id": "sub_91", "exercise_id": "ex_rate_limiter", "code": "..."})
monkeypatch.setattr("requests.get", fake_get)
assert fetch_submission("sub_91").exercise_id == "ex_rate_limiter"
That automatic revert is really the whole point of reaching for this fixture. Rolling your own version — setting os.environ["GRADER_CASE_TIMEOUT_SECONDS"] = "5" and resetting it by hand in a finally block — is deceptively easy to get wrong, and it fails in about the worst way possible: a test that dies partway through skips its own cleanup entirely, leaves the environment in a modified state, and now some unrelated test that runs afterward starts failing for reasons that have nothing to do with anything it actually does. monkeypatch guarantees the revert happens no matter how the test itself ends.
Patch where a name is looked up, not where it is defined
Nothing trips people up in Python mocking quite like this one, and seniority doesn't protect you from it — the mechanism behind it is completely mechanical once you see it, which makes it worth actually learning rather than re-discovering the hard way every few months.
Look at the top of the grader again:
# grader/storage.py
def fetch_submission(submission_id):
payload = requests.get(f"{API_URL}/submissions/{submission_id}", timeout=5).json()
return Submission(**payload)
# grader/service.py
from grader.storage import fetch_submission # grader.service.fetch_submission is now
# its OWN name, bound at import time
def grade(submission_id):
submission = fetch_submission(submission_id)
...
That from ... import did not create a reference to grader.storage. It copied the function object into a new name in grader.service's own namespace. There are now two names pointing at one function, and grade only ever reads one of them.
from unittest.mock import patch
# WRONG — patches grader.storage.fetch_submission, but grader.service
# never looks in grader.storage again after import
def test_grade_wrong():
with patch("grader.storage.fetch_submission", return_value=FIXED_SUBMISSION):
report = grade("sub_91") # still makes the real HTTP call
# RIGHT — patch the name where it is actually looked up
def test_grade_correct():
with patch("grader.service.fetch_submission", return_value=FIXED_SUBMISSION):
report = grade("sub_91") # the patch takes effect
Here's the rule, worth memorizing exactly as stated: you patch wherever the name gets looked up at call time, never wherever it happens to be defined. Because service.py imported the function directly by name, the patch target is grader.service.fetch_submission. Had service.py instead written import grader.storage and reached the function as grader.storage.fetch_submission() wherever it's invoked, then a patch aimed at grader.storage.fetch_submission would have worked just fine — because in that alternate version, the module goes and looks the name up fresh on grader.storage every single time, rather than holding its own private copy of the binding from import time.
The failure mode is worth recognizing on sight, because it does not look like a mocking problem. Your test hangs for eight seconds, or dies with a connection error, or — worst of all — silently succeeds against real data on a machine where the API happens to be reachable. If a patch appears to have no effect, the first question is always which module the code under test is actually reading the name from.