Skip to main content
CodeOath
← All posts

Testing64 min total · 12 parts

Testing Fundamentals: Unit, Integration, and E2E Tests Done Right

Part 4 of 12 · ~6 min

Test-Driven Development

Test-driven development turns the usual sequence on its head: instead of writing code and checking it afterward, you start with a test for behavior your codebase doesn't have yet, get it passing with the smallest change that will do, and only then worry about making the result presentable. The loop goes by a short name — red, green, refactor — and it runs over and over, in steps small enough that each one is almost boring.

  1. Red — write the test for the next sliver of behavior before any code exists to satisfy it. Run it and watch it fail; if it doesn't fail, either that behavior was already there or the test isn't actually checking anything.
  2. Green — write whatever gets that one test passing, and nothing more. Not the version you'd be proud to show a colleague, not the general-purpose solution. Just enough code to turn the red line green.
  3. Refactor — with a passing test now standing guard, this is the moment to actually clean things up: better names, less duplication, a simpler shape. Re-run it after every single change you make — green needs to hold the entire way through.

Walking the cycle: submissions the candidate is not allowed to make

A real requirement arrives. Right now, the submit endpoint stores whatever it is handed. If a candidate's window closed while they were still typing, their submission is accepted anyway — which is the production version of the > bug from the first chapter. The requirement: a submission that arrives outside the open window must be rejected, loudly, instead of silently stored.

Red — write the test first, against an AssessmentSession that does not exist yet:

test("a submission after the window has closed is rejected", () => {
  const session = new AssessmentSession(makeInvite({ closesAt: NOW - MINUTE }));
  expect(() => session.submit("def solve(): pass", NOW))
    .toThrow("Assessment window has closed");
});

Run it and it blows up immediately, since AssessmentSession doesn't exist yet — that's expected, and it's the red this whole cycle takes its name from. It's not just ceremony, either: this is the only moment you get actual proof the test is capable of failing. A test that has never once gone red is a test you have zero evidence about.

Green — write only as much code as it takes to turn that test green:

class AssessmentSession {
  constructor(invite) {
    this.invite = invite;
  }
  submit(code, now) {
    if (now >= this.invite.closesAt) throw new Error("Assessment window has closed");
  }
}

This is deliberately incomplete. It does not store the submission, it does not know about the candidate's ninety minutes, it does not handle an invite that has not opened yet. Its only responsibility at this stage is making the single test that currently exists turn green.

Red again — there's a case the last step left uncovered, so write the test for it before touching the implementation:

test("an on-time submission is recorded as an attempt", () => {
  const session = new AssessmentSession(makeInvite({ closesAt: NOW + HOUR }));
  session.submit("def solve(): pass", NOW);
  expect(session.attempts).toEqual([{ code: "def solve(): pass", at: NOW }]);
});

Green again — stretch the implementation just far enough to cover this new case too:

class AssessmentSession {
  attempts = [];
  constructor(invite) {
    this.invite = invite;
  }
  submit(code, now) {
    if (now >= this.invite.closesAt) throw new Error("Assessment window has closed");
    this.attempts.push({ code, at: now });
  }
}

Refactor — and this is the step people skip, so let us actually use it. Look at that guard. now >= this.invite.closesAt is a second, independent copy of a rule we already own. inviteStatus knows about the recruiter's window, the candidate's ninety minutes, invites that have not opened yet, and invites already submitted. This guard knows about one of those four things. Two copies of a rule drift apart, always, and the copy that drifts is the one nobody remembers exists.

submit(code, now) {
  // one rule, one place — and now a not-yet-open or already-submitted
  // invite is rejected too, which the hand-rolled guard never noticed
  if (inviteStatus(this.invite, now) !== "open") {
    throw new Error("Assessment window has closed");
  }
  this.attempts.push({ code, at: now });
}

Run both tests. Still green. That is the whole point of the safety net: this refactor genuinely changed behavior for three cases nobody had written a test for yet, and the two tests that do exist proved the cases they cover were not damaged on the way. Without them you would be squinting at the diff and hoping.

(The error message is now slightly wrong for the not-yet-open case, which is the kind of thing the next red step exists for.)

Where TDD genuinely helps, and where it is dogma

TDD pays off best on logic where the inputs and the correct outputs are both knowable in advance — inviteStatus fits that description exactly. Writing the assertion before the implementation forces you to actually commit to the contract instead of backing into one by accident. Type expect(inviteStatus(invite, NOW)).toBe("expired") for an invite whose recruiter window is still open but whose ninety minutes ran out, and you're forced to decide, on the spot, that the candidate's own clock wins — rather than finding out six weeks from now that nobody ever decided and the code simply picked a winner by accident. The same logic holds for parsers, validation rules, pricing math, state machines: anywhere the contract can be pinned down ahead of time. And there's a bonus nobody has to schedule separately — by the time you're done, you already own a regression suite, because writing it first was how you built the thing in the first place.

Follow it as a rule for the rule's own sake, though, rather than for what it actually buys you, and it turns counterproductive fast:

  • Exploratory or UI-heavy work is a poor match for the cycle, because you genuinely don't know yet what shape the solution should take. Our countdown banner is the case in point: four separate versions got thrown out over arguments about whether to show seconds once under five minutes remained. Writing a fresh test against each throwaway version isn't discipline, it's just rewriting tests exactly as fast as the design underneath them keeps changing.
  • Testing framework internals or a trivial getter just because some policy insists every line gets a test first results in tests written to placate the policy and nothing else. expect(session.invite).toBe(invite) is upkeep with no payoff — there's nothing real it could ever catch.
  • Writing at too fine a grain — a test per private method instead of a test per thing an outside caller can observe — lands you right back in the implementation-coupled trap from two chapters ago. TDD, on its own, does nothing to steer you away from that mistake. All it tells you is when to reach for a test, never what that test should actually be watching.

Here's the honest version: TDD is one good way to arrive at a clean design while a safety net catches mistakes along the way — it was never a moral requirement to write every test before its code, whatever any particular team's style guide claims. Plenty of genuinely well-tested systems were built test-after, and plenty of teams that followed the cycle religiously, step by step, still ended up with a bloated pile of tests that nobody six months later can explain the point of.