Testing64 min total · 12 parts
Testing Fundamentals: Unit, Integration, and E2E Tests Done Right
Part 6 of 12 · ~6 min
Mocking and Spies in Jest
Our system just grew its second service. When a candidate submits, the Node app has to hand the code to the Python grader, wait for a report, and email the recruiter:
// lib/assessment.js
import { findByToken, saveSubmission } from "./inviteRepo";
import { requestGrade } from "./graderClient"; // HTTP to the Python grader
import { sendScoreEmail } from "./notify"; // HTTP to an email provider
import { inviteStatus } from "./invites";
export async function submitAssessment(token, code, now) {
const invite = await findByToken(token);
if (inviteStatus(invite, now) !== "open") {
throw new Error("Assessment window has closed");
}
const submission = await saveSubmission(invite.token, code, now);
const report = await requestGrade(submission.id);
await sendScoreEmail(invite.recruiterEmail, report);
return report;
}
Eight lines, and testing it the naive way would spin up a container, execute a stranger's code, and send a real email to a real recruiter every time you saved a file. This is what mocking is for.
The vocabulary first, because people use these words loosely and it causes real confusion. A spy sits on top of a function that still does its real job — it just keeps a record of how that function got called, without changing what comes out of it. A mock goes further: it swaps the function or module out entirely for a stand-in whose return value, behavior, and failure modes you dictate yourself. You'll also run into stub for a stand-in that hands back a canned answer and keeps no record of anything, and fake for a working but stripped-down implementation — an in-memory grader, say, that returns a believable report without actually running a line of submitted code. Jest doesn't enforce these labels anywhere in its own API; what matters is knowing which behavior you actually need before reaching for jest.fn().
// jest.fn() creates a bare mock function — records calls, returns undefined unless configured
const sendScoreEmail = jest.fn();
sendScoreEmail("recruiter@example.com", { score: 80 });
expect(sendScoreEmail).toHaveBeenCalledWith("recruiter@example.com", { score: 80 });
expect(sendScoreEmail).toHaveBeenCalledTimes(1);
// Configuring what a mock returns
const loadExercise = jest.fn().mockReturnValue({ id: "ex_rate_limiter", caseCount: 10 });
const requestGrade = jest
.fn()
.mockResolvedValueOnce({ score: 80 }) // first call only
.mockRejectedValue(new Error("grader busy")); // every call after that
// jest.spyOn wraps an EXISTING method, so you can still call through to the real one
const tokenSpy = jest.spyOn(crypto, "randomUUID").mockReturnValue("iv_fixed");
// ... assert on a generated invite token without it changing every run ...
tokenSpy.mockRestore(); // put the real randomUUID back — important for isolation
That mockRestore is not a nicety. jest.spyOn mutates a real object in place, and a spy left installed leaks into every test that runs after it in the same file, which produces failures in tests that never mentioned crypto at all. Either restore it explicitly or put jest.restoreAllMocks() in an afterEach and stop thinking about it.
Mocking a whole module
Three of submitAssessment's four dependencies reach outside the process: the database, the grader, the email provider. The fourth, inviteStatus, is ours, pure, and instant. Hold onto that asymmetry — the next section is entirely about it.
// lib/graderClient.js
export async function requestGrade(submissionId) {
const res = await fetch(`${process.env.GRADER_URL}/grade`, {
method: "POST",
body: JSON.stringify({ submissionId }),
});
return res.json();
}
// lib/assessment.test.js
jest.mock("./inviteRepo"); // the database
jest.mock("./graderClient"); // the Python grader, over HTTP
jest.mock("./notify"); // the email provider, over HTTP
import { findByToken, saveSubmission } from "./inviteRepo";
import { requestGrade } from "./graderClient";
import { sendScoreEmail } from "./notify";
import { submitAssessment } from "./assessment";
// note: ./invites is NOT mocked — the real inviteStatus runs
beforeEach(() => {
findByToken.mockResolvedValue(makeInvite());
saveSubmission.mockResolvedValue({ id: "sub_91" });
});
test("a successful submission emails the recruiter the grader's report", async () => {
// Arrange — pin the one boundary whose answer this test is about
const report = { passed: 8, total: 10, score: 80, verdict: "pass" };
requestGrade.mockResolvedValue(report);
// Act
const result = await submitAssessment("iv_7c2a", "def solve(): pass", NOW);
// Assert — on the outcome, and on the one side effect that matters
expect(result).toEqual(report);
expect(requestGrade).toHaveBeenCalledWith("sub_91");
expect(sendScoreEmail).toHaveBeenCalledWith("recruiter@example.com", report);
});
jest.mock("./graderClient") swaps the entire module out for auto-generated mock functions across this whole test file — exactly the behavior we want here: no container spins up, no submitted code actually executes, no recruiter's inbox gets touched, and the whole test finishes in under a millisecond. When the auto-generated version isn't enough — say you need a stateful fake grader that returns a different report depending on which submission came in — Jest lets you write your own: park a hand-written implementation in __mocks__/graderClient.js, sitting alongside the real module, and jest.mock("./graderClient") picks it up automatically instead of generating a blank stand-in.
When mocking is right, and when the test stops testing anything
The whole point of mocking is to sever a test's connection to something genuinely outside what that test can or should control — a real call over the network, today's date, a payment processor, a slow database, a whole container runtime spinning up. Used for that, it's the right call, and requestGrade is about as textbook a case as you'll find.
It stops paying for itself the moment you mock the exact thing the test is supposed to be checking. Look at what happens here:
// This test provides almost no signal.
// It mocks inviteStatus, then asserts that inviteStatus was called.
jest.mock("./invites");
import { inviteStatus } from "./invites";
test("submitAssessment checks the invite status (weak)", async () => {
inviteStatus.mockReturnValue("open"); // ← we have to supply the answer ourselves
await submitAssessment("iv_7c2a", "def solve(): pass", NOW);
expect(inviteStatus).toHaveBeenCalled();
});
The name says it checks the invite status. What it actually checks is that a function it replaced with a stand-in got called. inviteStatus — the function we spent two chapters getting right, with its five branches and two competing clocks — never runs. Break every one of those branches, ship an inviteStatus that returns "open" for an invite that closed last Tuesday, and this test stays green forever, because the real one was swapped out before any logic could run.
The marked line is the tell, and it is worth learning to flinch at. The test had to be told what the invite's status is, because the thing that knows how to work that out is exactly what the test removed. Any time you find yourself hand-feeding a mock the answer that your own code was supposed to compute, you have mocked one layer too deep.
A useful version does one of two things. Either let the real inviteStatus run — it is pure, instant, and has no dependencies, so there is no reason on earth to fake it — and assert on the outcome:
test("submitAssessment refuses a submission after the window closed", async () => {
findByToken.mockResolvedValue(makeInvite({ closesAt: NOW - MINUTE }));
await expect(submitAssessment("iv_7c2a", "def solve(): pass", NOW))
.rejects.toThrow("Assessment window has closed");
expect(requestGrade).not.toHaveBeenCalled(); // nothing was sent to the grader
});
Or, if you only care about the rule itself, drop a layer and test inviteStatus directly, where you can cover all five branches for the price of one test.each table. That second option is usually the better one.
A rule that holds up in practice: save mocking for what's genuinely outside your control or reliably slow — someone else's API, the system clock, the filesystem, another team's service entirely. Stay much more reluctant when the target is your own business logic, reached for only because it would make a test run faster or feel more isolated. Wanting to usually means one of two things is true: either the function under test has taken on too many responsibilities, or this test doesn't belong at this layer at all and should be written one level down, pointed directly at the logic it's actually trying to check.