Testing64 min total · 12 parts
Testing Fundamentals: Unit, Integration, and E2E Tests Done Right
Part 12 of 12 · ~7 min
Common Testing Anti-Patterns
Certain habits keep resurfacing in suites that were perfectly healthy once and quietly turned into a liability over time. All three are already somewhere in what we have built.
Testing implementation details instead of behavior
We already spent a full chapter on this, but it deserves to stand on its own in a list of anti-patterns too. Any test that checks how an answer got computed — which private helper fired, how many times, what shape some intermediate value happened to take — is one reorganization away from breaking, regardless of whether anything a caller could actually see changed at all. Our expect(spy).toHaveBeenCalledTimes(1) test on candidateDeadline is the specimen sitting right here in this system: fold that two-line helper back into lib/invites.js, a change no user could ever notice, and that one test goes red while every behavior-focused test stays exactly where it was.
Ask around and this turns out to be the number-one reason a team ends up describing its own suite as something that fights every refactor instead of making refactors safer — and the damage snowballs, because once cleanup gets punished, people simply stop doing it. The fix hasn't changed since the chapter where this first came up: assert on what a caller actually depends on going in and coming out, never on machinery the caller can't see and has no business caring about.
Flaky tests
A flaky test is one that alternates between passing and failing with no actual code change happening in between. What it really costs a team isn't the occasional red build — it's the erosion of trust that follows. The instant a team decides some test "just does that sometimes," the habit that forms is re-running CI on red instead of reading why it failed, and that habit doesn't stay contained to the flaky test — the next genuine failure from that same test ends up dismissed the exact same way. A single unreliable test can quietly drag the whole suite's credibility down with it.
There are four usual root causes, and every one of them has already appeared in this reference:
- Time dependence. The very first test in the first chapter, calling
Date.now()three times and betting the gap stays under sixty seconds. Its close cousin is the hardcoded date: a test written in March withclosesAt: 1_726_300_000_000because that was comfortably in the future at the time, which starts failing without warning on a Saturday in September. And the end-to-end variant: an assertion that the grader's report appears within two seconds, which is true on a quiet machine and false the moment CI is running four jobs at once. - Order dependence. The
workspacefixture widened tosessionscope for speed, so one test'ssolution.pyis still sitting there for the next one. Or theafterEachthat truncates the invites table getting deleted because "nothing seemed to depend on it." The tell is unmistakable: the test passes in the full suite and fails when you run it alone, or vice versa. - Real network or real external services. The grader test that pulls its sandbox image from a real registry inherits every bit of that registry's slowness, throttling, and downtime as intermittent failures of its own. So does any test that reaches a shared staging environment, which additionally means your suite's result depends on what a colleague happened to deploy there four minutes ago.
- Unawaited async work. The exact bug from the async chapter: a test that reports success purely because it wrapped up before its own assertion ever got the chance to run. This one is the most treacherous of the four, because it usually presents as passing, and only flips to a real failure when timing shifts slightly — a faster CI machine, a mocked call that now resolves on a different tick — at which point it looks like a brand-new failure in code nobody touched.
Each cause has its own specific remedy, and we've already built every one of them into this system: a clock passed in as an argument and constants pinned to a frozen NOW for the time problem; function-scoped fixtures with a real afterEach for the ordering problem; mocking the actual external boundary for the network problem; await or return for the async problem. What never changes is the first move, and it's a judgment call rather than a technique: don't wave off an intermittent failure as nothing. Odds are overwhelming it's one of these four, and odds are just as good it's fixable before lunch — a far better deal than watching a team gradually stop believing anything their own suite tells them.
Over-mocking
This one gets a name of its own because it creeps in gradually and is genuinely hard to notice from inside the suite you're writing. Once nearly every dependency in a test file has been swapped for a mock, what's actually being verified is that your code dialed its mocks in the right sequence — nothing more than that. A real regression can walk straight through a suite like that without tripping anything, because the real logic was never in the room. Only a stand-in someone configured by hand was.
We wrote the specimen back in the Jest mocking chapter and then used it again as a coverage-gaming example, which is not a coincidence: the two anti-patterns produce the same test. jest.mock("./invites") replaced five branches of real business logic with a stand-in, the test hand-fed that stand-in the answer, and then it asserted that the stand-in had been called.
You can spot it by pattern: every assertion reads expect(mockThing).toHaveBeenCalledWith(...), and not one of them ever touches a real computed value or a genuinely observable outcome. Catch a suite drifting that direction and walk it dependency by dependency, asking the same question of each mock in turn — is this standing in for something genuinely outside the test's control, like the grader, the email provider, the clock, the container runtime, or is it quietly burying the one piece of logic this test claimed to be checking? requestGrade answers that question cleanly. inviteStatus never had any business being asked it.
Trace all three anti-patterns back far enough and they share one origin: forgetting what a test is actually there to do. The entire point of writing one is confidence — proof that real behavior works right now and will keep working later. Not a number on a coverage dashboard. Not a green checkmark in CI. Not compliance with a rule that says tests come first.
That is worth checking against the system we built. There is now a test.each table pinning all five branches of inviteStatus, an integration test that round-trips a real invite through a real database, a parametrized table over the grader's five exit codes, a fixture that destroys its sandbox even when the test explodes, and one end-to-end test that follows a candidate from an emailed link to a submitted solution. Between them, none of the bugs in this reference survives: not the > that let a late submission through, not the missing submittedAt in the row mapping, not the submissionId/submission_id mismatch, not the ZeroDivisionError on an exercise with no graded cases, and not the hardcoded timestamp that quietly went stale.
The one test that would survive all of them is the one that mocked inviteStatus and asserted that it was called — which is the whole lesson, in one line. Every practice here, the pyramid and AAA and TDD and careful mocking and coverage read as a negative signal, points back at the same question from a different angle: if this thing were broken, would this test tell me?
Want to build these reflexes rather than just read about them? The code lab has runnable exercises in both Jest and pytest. If the language fundamentals underneath any of this feel shaky, Python Fundamentals for Interviews and JavaScript Core Concepts cover the layer these testing ideas sit on top of, and CI/CD Pipelines Explained picks up right where this leaves off — how a suite like the one we just built actually gets triggered automatically on every push.
Continue learning
- Interview & Career PrepThe Non-Technical Half of the Interview: Behavioral Questions, the STAR Method, and What Recruiters Are Actually Scoring
- AI & LLM EngineeringAI & LLM Engineering Fundamentals: Prompting, RAG, Embeddings, and Function Calling
- TypeScriptTypeScript Fundamentals: Types, Interfaces, Generics, and Why It Catches Bugs Before Runtime