Testing64 min total · 12 parts
Testing Fundamentals: Unit, Integration, and E2E Tests Done Right
Part 2 of 12 · ~6 min
Why We Test
Strip it down and a test is nothing more than code whose entire job is running other code and checking the result itself — the same check a person would make by clicking around the product, except it happens on its own, the same way, as many times as anyone wants, with nobody sitting at a keyboard for it.
That last part is where the actual payoff lives, and it is easier to see against numbers than as a slogan. Checking inviteStatus by hand means finding or faking an invite that has not opened yet, one that is open, one that closed yesterday, and one already submitted — four invites, four rows in a database, four clicks. Then the recruiter asks for a feature: the candidate's own ninety minutes should run out independently of the recruiter's window, whichever comes first. Now there are two clocks instead of one, so what was a four-case hand-check becomes closer to a dozen, and someone has to redo all twelve by hand every time this file changes, forever. This is exactly why doing it by hand stops working once a product has more than a couple of moving parts: every feature you add is one more thing a person has to remember to click through before anything can ship, and that list only ever gets longer.
The testing pyramid
Tests come in different shapes, and they cost wildly different amounts to run — which is why a suite that stays pleasant to work with skews hard toward the cheap end instead of spreading evenly across all three. Here is that shape, with our own system's tests dropped into each layer:
▲
/ \ End-to-End (few)
/---\ — open the emailed link, type, submit, see the score
/ \
/ Integ.\ Integration (some)
/---------\ — findByToken against a real Postgres; Node calling the real grader
/ \
/ Unit \ Unit (many)
/-----------------\ — inviteStatus, remainingMs, classify_case, percentage_score
- Unit tests stay inside one function, one class, one module, with everything around it faked or removed — this is the layer
inviteStatusitself belongs to. Feed it an object and a number, read a string back, and know instantly which function is wrong when the answer doesn't match. Nothing real runs underneath it: no live database, no socket opened, no file touched. You can afford thousands of these because each finishes in a few milliseconds and points almost straight at the broken line. - Integration tests put two or more real pieces in the same room at once — our code against a genuine, disposable Postgres instance, say, or two of our own modules calling each other for real instead of through a stand-in one side made up. There are deliberately fewer of these, each one takes noticeably longer, and each one is checking more surface area in a single run.
- End-to-end tests replay the whole thing the way an actual candidate experiences it: a real browser opens the emailed link, types into the real editor, clicks the real submit button, against a backend that is genuinely running. These are the rarest of the three, by far the slowest, and closest to the truth of what a user will actually see happen.
Why lean this hard toward the cheap layer? Every test buys some amount of confidence and costs some amount of time, and that exchange rate gets worse the higher up the pyramid you go, right as your budget for how many you can actually run keeps shrinking. Two thousand unit tests can finish before your fingers leave the keyboard; a hundred integration tests fit comfortably on every push; a dozen end-to-end tests before a deploy is still tolerable — stack those three together and you get a suite people run constantly without resenting it. Invert the shape instead — heavy on end-to-end, thin on unit — and the same amount of testing now takes the better part of an hour, turns red for reasons disconnected from the actual defect roughly once in every three runs, and gets quietly skipped the moment a deadline gets tight, precisely because it costs so much to run and tells you so little when it fails.
Common mistake: assuming more end-to-end coverage is automatically more thorough, and therefore automatically better. It is not automatically either. End-to-end tests genuinely catch classes of bug — rendering problems, real cross-service failures — that nothing running inside a unit test's sandbox could ever see. But that same realism is exactly what makes the layer slow, prone to failing for unrelated reasons, and miserable to debug: when one goes red, the cause could be your code, a race condition, a stale row left in the test database, or the test harness itself having a bad day. The pyramid is shaped this way because the confidence you get per second of CI time falls off sharply the higher you climb — not because the top layers are somehow worse tests, just far more expensive ones.
The cost curve of a bug, by where it is caught
Take one specific defect and follow it. Look again at the third comparison in our starting version:
if (now > invite.closesAt) return "expired";
That should be >=. At the exact millisecond the window closes, this says "open". It is one character, and how much it costs depends entirely on how far it travels before anyone notices:
| Where the bug is caught | Typical cost to fix | Why |
|---|---|---|
| Editor / while writing the code | Seconds | You still hold the whole context in your head; nothing has been committed |
| A failing unit test, pre-commit | Minutes | Still local, still fresh context, nobody else has touched the code |
| Code review | Tens of minutes | A second person has to re-derive the context; a round trip of comments and re-pushes |
| CI on a pull request (integration/end-to-end) | An hour or more | Full pipeline re-run, possibly blocking other people's merges |
| Production | Hours to days, plus damage | Incident response, rollback or hotfix, customer impact, a post-mortem — and the author has mentally moved on |
The bottom row is not hypothetical for a bug like this one. In production it means a candidate submits a fraction of a second after their window closed, the submission is accepted, the grader scores it, and a recruiter makes a hiring decision on work that was supposed to be out of time. Now you are not fixing a comparison operator, you are deciding what to do about a score that already went out.
That's really the whole case for catching things early: a production bug isn't costly because it's embarrassing, it's costly because fixing the exact same defect gets more expensive purely as a function of how long it sat there — by the time anyone notices, other code has been written on the assumption it was fine, and whoever eventually touches it has to rebuild context someone else already had for free while writing the original line. Ten seconds spent on a unit test that catches this the day it's introduced is about as good a trade as an engineer ever gets offered.