Skip to main content
CodeOath
← All posts

Testing64 min total · 12 parts

Testing Fundamentals: Unit, Integration, and E2E Tests Done Right

Part 11 of 12 · ~5 min

Code Coverage

Code coverage is a measurement of how much of your codebase actually got exercised while the suite ran. The two numbers you'll see quoted are line coverage, the share of executable lines that ran at least once, and branch coverage, the share of decision points — every if, every else, every fork — that got taken in both directions, not just one.

The grader has a function that turns a score into a hiring signal:

# grader/scoring.py
PASS_MARK = 70

def verdict(score):
    if score >= PASS_MARK:
        return "pass"
    return "review"

def test_verdict_passes_a_strong_score():
    assert verdict(85) == "pass"

That single test gives verdict 75% line coverage — three of its four statements ran; return "review" never did — and 50% branch coverage, because the if was only ever taken in one direction. Both numbers are honest here, and they roughly agree.

Now watch them come apart. This is the line the recruiter's email leads with:

def score_summary(report):
    line = f"{report['passed']}/{report['total']} cases passed"
    if report["timeouts"]:
        line += f" ({report['timeouts']} timed out)"
    return line

def test_score_summary_mentions_timeouts():
    assert score_summary({"passed": 8, "total": 10, "timeouts": 2}) == "8/10 cases passed (2 timed out)"

One test, and now every executable line in score_summary has run: 100% line coverage. Branch coverage is still 50%, because the if has never been false. The untested path is the ordinary one — a submission with no timeouts at all — and it is the path where a stray edit turning line += into an unconditional append would put "(0 timed out)" in every recruiter's inbox, with the coverage report at a clean hundred percent throughout.

That's the asymmetry worth carrying away. Line coverage gets quoted far more often, and it's also the easier of the two numbers to game, because a line that merely executed tells you nothing about whether every path worth caring about inside it actually got exercised.

That gap is where our very first bug lived. Go back to the opening version of inviteStatus:

if (now > invite.closesAt) return "expired";   // should be >=

A test suite can hit 100% line coverage on that function with four tests, one per returned status, and still never test the single instant the bug lives at. Coverage tools count the lines that ran, and that line ran — four times. The boundary is not a line. It is a value.

100% coverage, still a real bug

What coverage actually measures is execution, never correctness — a line running to completion, every single time, tells you nothing about whether it computed the right thing, so long as nobody ever checks what it handed back:

def percentage_score(passed, total):
    return round(100 * passed / total)   # no guard against total == 0

def test_percentage_score():
    assert percentage_score(8, 10) == 80   # covers the one line — 100% line coverage

percentage_score is now at 100% line coverage, and it has a live crash in it. When a recruiter publishes an exercise and forgets to add any graded cases — which happens, because the form lets you — load_cases returns an empty list, total is 0, and the grader raises ZeroDivisionError in the middle of processing a real candidate's real submission. The coverage report will never flag it, because the report only ever asks whether a line executed, never whether the executed behavior was correct for every input that matters.

The test that catches it is the one that asks a question the coverage number cannot:

def test_percentage_score_handles_an_exercise_with_no_cases():
    assert percentage_score(0, 0) == 0   # fails today: ZeroDivisionError

Nothing about coverage got you there. Thinking about the inputs did.

Why chasing 100% is usually the wrong goal

Think of coverage as a warning light, not a gold star. When it's low, that's a trustworthy signal — something in there genuinely has zero eyes on it, which is worth knowing. When it's high, that tells you far less than it sounds like it should: all it confirms is that the code ran at some point during the suite, which — as percentage_score just showed — is a much weaker claim than "the code works."

Chase the number itself as the goal and you get a predictable outcome: tests engineered to touch a line, assertion-free or nearly so, existing purely to nudge a percentage upward rather than to check that anything actually works. Our own suite already has a specimen of exactly this, and we wrote it several chapters back:

jest.mock("./invites");
test("submitAssessment checks the invite status (weak)", async () => {
  inviteStatus.mockReturnValue("open");
  await submitAssessment("iv_7c2a", "def solve(): pass", NOW);
  expect(inviteStatus).toHaveBeenCalled();
});

That test runs through every line submitAssessment has. It pushes the coverage number up. And there is no way for it to ever fail for a reason connected to whether the system actually does its job. That's worse than skipping a test for that function entirely, because a green line on a dashboard now claims a safety that doesn't exist — and worse, it discourages anyone from ever writing the real test, since as far as the tooling is concerned this one is already done.

Point your effort instead at where a bug would genuinely cost something. inviteStatus decides whether someone's work counts at all, percentage_score decides what number a recruiter actually reads — both deserve every branch and every edge nailed down. A __repr__ method on the submission object does not deserve the same attention. Watch for coverage sagging on a path that matters; don't lose sleep over a report that reads 94% instead of 100% purely because a few one-line wrappers nobody depends on never got a test written for them.