Skip to main content
CodeOath
← All posts

Testing64 min total · 12 parts

Testing Fundamentals: Unit, Integration, and E2E Tests Done Right

Part 3 of 12 · ~10 min

What Makes a Good Unit Test

Having a pile of unit tests doesn't automatically mean you have a good pile of unit tests. Five properties separate the ones worth trusting from the ones that just take up CI time, and remembering all five gets a lot easier once you notice their initials spell a word: Fast, Isolated, Repeatable, Self-checking, Timely — FIRST. Nodding along to a list like that costs nothing, so instead let's hold the one test we've actually written up against each letter and see where it comes apart.

Here is the only test that exists so far:

test("an open invite is open", () => {
  const invite = {
    opensAt: Date.now() - 60_000,
    closesAt: Date.now() + 60_000,
    submittedAt: null,
  };
  expect(inviteStatus(invite)).toBe("open");
});

It is green. It is also failing two of the five properties, and one of those failures is going to haunt us for the rest of this reference.

  • Fast. Ours finishes in a few milliseconds, which is the bar to clear. Anything in a unit suite that takes ten minutes to run has stopped being a unit test somewhere along the way — it's almost certainly touching a live database or making a real call over the network.
  • Isolated (independent). Nothing about this test's outcome should hinge on some other test having already run, on whatever order the runner happens to pick, or on state left lying around from earlier. Run it by itself, run it after fifty others, run it a hundred times back to back — the answer should not move. Ours qualifies: it builds a brand-new invite object out of nothing and never reaches for anything shared.
  • Repeatable. ❌ Run this against the same code a hundred times and the answer should not budge — not "usually," not "passes on my machine but not CI." This test calls Date.now() three times — twice in the setup, once inside inviteStatus — and those three calls return three different numbers. Today that gap is a millisecond and the test passes. Under CI load, with a garbage collection pause landing in the wrong place, it is a variable gap, and the whole test is an implicit bet that no more than sixty seconds elapse between two lines. That bet is extremely likely to pay off, which is exactly what makes it dangerous: it will pay off a thousand times and then fail once, at 2am, on someone else's unrelated pull request.
  • Self-checking. Pass or fail falls out of the assertion itself — nobody has to squint at a console log and decide whether a number looks about right. Something that only prints a value for a human to eyeball isn't really a test, it's a print statement wearing a costume. Ours clears this one too.
  • Timely. ❌ The test should show up around the same time as the code it covers, ideally slightly before, back while you still remember exactly why you're writing it rather than reconstructing an approximation of that reason three sprints later. This one got bolted on afterward, which is exactly why it only bothers checking the one status that happened to be easiest to set up.

That is two clear failures and, underneath them, a third problem that caused both: the function reaches out and grabs the clock itself.

function inviteStatus(invite) {
  const now = Date.now(); // ← the source of the trouble
  // ...the rest of the function, unchanged
}

There is no value you can pass in to make inviteStatus believe it is any particular moment, so every test has to construct its invite relative to right now and hope the wind does not change mid-test. Testing the "expired" path means building an invite that closed in the past; testing a ninety-minute clock that has fifteen minutes left means arithmetic against a moving target.

The fix is one parameter:

// lib/invites.js — version 2: the clock comes in from outside
function inviteStatus(invite, now) {
  if (invite.submittedAt !== null) return "submitted";
  if (now < invite.opensAt) return "early";
  if (now >= invite.closesAt) return "expired";
  return "open";
}

Date.now() still gets called — just one layer up, at the edge of the system, by whoever is handling the actual request. Everything below that edge takes the moment as an argument. That single change is the difference between a function you can only observe and a function you can interrogate, and it is worth naming as a general rule: anything a function reaches out and takes for itself — the clock, a random number, an environment variable, the network — is something a test cannot control. Passed in as an argument, it becomes something a test can pin down exactly. We will hit this same idea again under a different name when we get to mocking, because it is the same idea.

(That version also quietly fixes the > to >=. Hold that thought — we come back to it in the coverage chapter, where it does some real work.)

Two small habits that make the rest of this possible

The clock now comes in from outside, so tests need something to pass it. Two pieces of test scaffolding get set up once here and then used by every test for the rest of this reference.

A frozen instant, and everything described relative to it.

// test/helpers.js
const MINUTE = 60_000;
const HOUR = 60 * MINUTE;
const DAY = 24 * HOUR;
const NOW = 1_726_300_000_000; // 2024-09-14T07:46:40Z — frozen, on purpose

Every invite in every test from here on is described as NOW - 30 * MINUTE or NOW + 2 * DAY, never as Date.now() + something. There is no clock in the tests at all, which means there is nothing to drift, which is the Repeatable property delivered rather than hoped for.

A builder for the object under test. makeInvite fills in every field with a sensible default and lets each test override only the ones it actually cares about:

// test/helpers.js
function makeInvite(overrides = {}) {
  return {
    token: "iv_7c2a",
    candidateEmail: "candidate@example.com",
    recruiterEmail: "recruiter@example.com",
    exerciseId: "ex_rate_limiter",
    opensAt: NOW - 1 * DAY,
    closesAt: NOW + 2 * DAY,
    durationMs: 90 * MINUTE,
    startedAt: null,
    submittedAt: null,
    ...overrides,
  };
}

This is not just less typing. A test that spells out all nine fields buries its own point: a reader has to compare nine values against nine other values to work out which one the test is actually about. makeInvite({ startedAt: NOW - 91 * MINUTE }) says the point out loud — this test is about a candidate whose ninety minutes ran out, and nothing else here matters. It also means that when the invite grows a tenth field next quarter, you change one file instead of four hundred tests.

Test behavior, not implementation

Here's the real design question buried inside every unit test: what is this test actually allowed to know about the code it checks?

A test that sticks to calling the public interface and inspecting what comes back survives a function being rewritten from scratch underneath it — the interface didn't move, so the test never notices. A test that pokes at private state, tallies how many times some internal helper fired, or pins down the exact order internal steps happened in is a different animal: it snaps the instant anyone reorganizes that code, whether or not anything a caller could actually observe changed at all.

Our system is about to give us a perfect example. The recruiter's feature request lands: the candidate's ninety minutes should run out independently of the recruiter's window, whichever ends first. So inviteStatus grows a helper, and the helper lands in a module of its own:

// lib/deadlines.js
export function candidateDeadline(invite) {
  return invite.startedAt === null ? Infinity : invite.startedAt + invite.durationMs;
}
// lib/invites.js — version 3: two clocks, whichever runs out first
import * as deadlines from "./deadlines";

export function inviteStatus(invite, now) {
  if (invite.submittedAt !== null) return "submitted";
  if (now < invite.opensAt) return "early";
  if (now >= invite.closesAt) return "expired";
  if (now >= deadlines.candidateDeadline(invite)) return "expired";
  return "open";
}

Now here are two tests for that new behavior. Both pass right now. Only one of them is worth having.

// Implementation-coupled — this test is watching the machinery, not the result.
// Inline candidateDeadline, rename it, or fold the two expiry checks into a
// single comparison, and this goes red while the behavior stays identical.
import * as deadlines from "./deadlines";

test("inviteStatus consults candidateDeadline (bad)", () => {
  const spy = jest.spyOn(deadlines, "candidateDeadline");
  inviteStatus(makeInvite({ startedAt: NOW - 10 * MINUTE }), NOW);
  expect(spy).toHaveBeenCalledTimes(1);
  spy.mockRestore();
});

// Behavior-focused — this test only cares about the one thing anyone outside
// this module can actually observe. Survives any internal rewrite.
test("an invite whose ninety minutes have run out is expired (good)", () => {
  const invite = makeInvite({
    startedAt: NOW - 91 * MINUTE,
    durationMs: 90 * MINUTE,
    closesAt: NOW + 2 * DAY, // the recruiter's window is still wide open
  });
  expect(inviteStatus(invite, NOW)).toBe("expired");
});

The second test is also a better bug report. When it fails, it tells you that a candidate who ran out of time is still being treated as active — something a human can act on. When the first one fails, it tells you a function you have never heard of was called a different number of times than someone expected two years ago.

Here's a quick test for the test itself: picture swapping out everything inside the function — a swapped-out loop, a rewritten helper, a wholly different algorithm — while the outward behavior stays identical down to the last detail. If your test would still break under that swap, it was never really checking behavior, it was checking a particular implementation you happened to have lying around at the time. That's why a codebase can be covered top to bottom and still be something people dread refactoring — the problem was never how many tests existed, it's what those tests were watching.

The AAA structure

Look at enough good unit tests and a pattern jumps out: nearly all of them break into the same three moves, commonly labeled Arrange-Act-Assert. Here's that shape applied to remainingMs, the function actually driving the countdown a candidate stares at while working:

// lib/invites.js
export function remainingMs(invite, now) {
  if (inviteStatus(invite, now) !== "open") return 0;
  return Math.min(invite.closesAt, deadlines.candidateDeadline(invite)) - now;
}
test("remainingMs counts down the candidate's clock when it ends first", () => {
  // Arrange — started 30 minutes ago on a 90-minute clock,
  // inside a recruiter window that stays open for another two days
  const invite = makeInvite({
    startedAt: NOW - 30 * MINUTE,
    durationMs: 90 * MINUTE,
    closesAt: NOW + 2 * DAY,
  });

  // Act — the one call actually under test
  const result = remainingMs(invite, NOW);

  // Assert — one behavior, one expectation
  expect(result).toBe(60 * MINUTE);
});

Keep those three parts visually distinct — one blank line between them is enough, a comment does it even better — and a test turns into something you can read at a glance: here's what went in, here's what ran, here's what ought to come out the other side. Tangle setup, the call, and the checks together and you get something nobody can scan quickly, which is usually the tell that the test is quietly checking three things instead of one. Aim for one test, one behavior, one reason it could ever go red — if you can't summarize what a test checks in a single plain sentence, odds are it's checking more than that.

One more thing worth noticing about that test: the Assert section has exactly one expect in it. That is not a hard rule — asserting on two fields of one returned object is fine — but it is a useful pressure. When a test accumulates six unrelated assertions, the first failure hides the other five, and the name at the top of the test has stopped describing what it does. Splitting it is almost always the right move.