Testing64 min total · 12 parts
Testing Fundamentals: Unit, Integration, and E2E Tests Done Right
Part 10 of 12 · ~5 min
Integration vs. End-to-End Testing
Every test so far has been a unit test, which means every one of them has been deliberately blind to the same category of bug: the kind that lives between two pieces rather than inside either one.
Our system has three real seams, and each has already produced a bug that every unit test in this reference would sail straight past.
Seam one: our code and the database. The invites table has a column named closes_at. The repository layer maps a row to an object, and somebody typed the mapping by hand:
function rowToInvite(row) {
return {
token: row.token,
candidateEmail: row.candidate_email,
recruiterEmail: row.recruiter_email,
exerciseId: row.exercise_id,
opensAt: row.opens_at,
closesAt: row.closes_at,
durationMs: row.duration_ms,
startedAt: row.started_at,
// submittedAt: forgotten
};
}
Every unit test of inviteStatus passes, because every one of them builds its invite with makeInvite, which always sets submittedAt. In production invite.submittedAt is undefined, undefined !== null is true, and every invite in the system reports as "submitted" the moment it is loaded from the database. Nobody can start an assessment. The function is correct; the object being handed to it is not. An integration test that writes a real row and reads it back through the real repository catches this in one assertion:
test("an invite round-trips through the database with every field intact", async () => {
await db.insertInvite(makeInvite({ token: "iv_7c2a" }));
const loaded = await findByToken("iv_7c2a");
expect(loaded).toEqual(makeInvite({ token: "iv_7c2a" }));
});
Seam two: our two services. The Node client posts { submissionId }. The Python grader's endpoint reads payload["submission_id"]. Both sides are impeccably unit-tested — the Node tests assert the client posts what the client posts, the Python tests call grade() with a submission id directly — and the two sides have never once been in the same process. Every submission in production fails with a KeyError. camelCase meeting snake_case across a language boundary is one of the most reliably recurring bugs in any polyglot system, and it is structurally invisible to unit tests on either side.
Seam three: our code and the database's rules. submissions.code is a varchar(10000). A candidate pastes a solution with a large embedded test harness — twelve thousand characters. Our validation, thoroughly unit tested, is happy. Postgres is not, and the insert raises. The constraint exists in the schema, and the schema is precisely what a unit test replaces with an object literal.
This is precisely the gap integration tests are built to close — put two or more real collaborators in the same test and check that the boundary between them actually holds. A query that parses fine but pulls the wrong rows once it hits a real schema, one side's serialized output that the other side can't parse, a database constraint quietly rejecting data your validation logic was sure was fine — none of these show up in a unit test, however thorough, because the bug isn't sitting inside either piece. It's living in the handshake between them.
End-to-end tests push a level further still, running the entire assembled system the way an actual candidate would sit down and use it. That calls for a real browser, so a different tool drives it — Playwright in our case, though Cypress or Selenium would play the same part. Here's a single test that follows one person's whole journey, start to finish:
test("a candidate can open an invite, write a solution, and submit it", async ({ page }) => {
await page.goto(`${BASE_URL}/a/iv_7c2a`); // the emailed link
await expect(page.getByRole("timer")).toContainText("1:30:00");
await page.getByRole("textbox", { name: "Editor" }).fill("def solve(n):\n return n * 2\n");
await page.getByRole("button", { name: "Submit" }).click();
await expect(page.getByText("Submission received")).toBeVisible();
});
That one test covers ground neither of the other layers can reach. It is the only thing in the whole suite that would notice a sticky footer rendering on top of the submit button, so that every unit test passes and no candidate on a laptop can physically click it. It is the only thing that would catch the editor's web worker failing to load because production's content-security-policy header blocks it. And it is the only thing that would catch GRADER_URL pointing at staging in the production config — a bug where every line of code is perfect and the system is still broken.
Climb this ladder and the same trade keeps repeating, rung after rung:
| Unit | Integration | End-to-end | |
|---|---|---|---|
| Speed | Milliseconds | Tens of ms to a few seconds | Seconds to minutes |
| Realism | Low — most collaborators are faked | Medium — some real, some faked | High — the real system, close to real usage |
| Flakiness risk | Very low | Low to moderate | Highest — timing, network, environment all in play |
| Failure locates the bug | Very precisely | Somewhat — narrows to the seam, not the line | Poorly — "something broke somewhere in this flow" |
| Typical count in a healthy suite | Hundreds to thousands | Dozens to low hundreds | A dozen to a few dozen, on critical paths |
Trace what happens at the bottom-right cell of that table, against the end-to-end test above. A red result there tells you only that a candidate couldn't submit — the actual cause could be hiding in the invite link, the timer component, the editor, the submit handler, the repository, the grader, the email provider, a CSP header, or nothing more than the test's own flaky timing. That vagueness is the toll you pay for realism, and it's exactly why the pyramid keeps its shape: each layer is watching for a genuinely different class of failure, so none of the three can stand in for either of the others. Getting to "100% unit tested" doesn't stop a team from shipping two services that are each individually flawless and structurally unable to talk to each other. The sane allocation stays the one from the first chapter: put the bulk of your effort into the cheap, fast, precise layer, and spend the pricier layers on purpose, reserved for the small number of things nothing else in the suite is positioned to catch — not as a substitute for the coverage the unit layer already gives you for free.