Testing — study guide

The concept's fragments, read in order.

Confidence that it still works

A test is a small program that checks another program does what it should. You describe a situation, run the code, and compare what came out against what you expected. Automated tests do this the same way every time, on demand, without a human clicking through the application by hand and remembering what to look for.

The reason to bother is change. Code that works today gets edited tomorrow, and an edit meant to fix one thing quietly breaks another. A suite of tests is like a safety net under a trapeze - it does not stop you from making a bold move, it catches you when one goes wrong: it does not stop you from changing the code, it catches you the moment a change breaks something that used to work. Running the suite after every edit turns "I think this still works" into "the checks say it still works."

The alternative is verifying by hand, which is fine for a program you touch once and slow torture for one you touch daily. Every manual check is a check someone has to remember to run and run correctly. The rest of this concept is the shapes those automated checks take: what a single test is made of, the kinds of test that cover a small piece versus the whole system, how many of each to write, and the discipline of writing the test before the code.

One small piece, in isolation

A unit test checks a single small piece of code on its own — typically one function, called directly with inputs you choose and checked against the output you expect. It has the narrowest scope of any test: nothing else is involved, so nothing else can be blamed. If the function is supposed to turn a price and a tax rate into a total, a unit test hands it a price and a rate and confirms the total is right.

Because a unit test touches so little, it is significantly faster than tests that reach outside the code, and a project accumulates many of them — often one or several per function. Fast and numerous is the point: a suite of thousands can run in seconds, so you can run it after every small edit rather than saving verification for the end.

The narrow scope also makes a failure precise. When one function's test fails, the defect is in that function, not somewhere across a chain of moving parts. A unit test does not tell you the whole system works, but it tells you exactly which small piece stopped working, which is often the first thing you want to know.

Do the pieces fit together

An integration test checks that separate components work together rather than each one alone. A function might pass its unit test perfectly and still fail the moment it has to save a record to a real database, or hand its output to another module that expected a slightly different shape. Integration tests exist to catch the gaps between parts — the seams a unit test never crosses because it deliberately stays inside one piece.

The wider scope has a cost. An integration test usually has to involve the real thing it integrates with — a database, a filesystem, a second service — and spinning that up and talking to it is slower than calling a function in isolation. These tests are also harder to write, because you have to set up the outside part and put it back the way you found it.

So integration tests earn their keep in the middle: fewer than the unit tests, slower to run, but checking something no unit test can. They answer a different question — not "is this function correct" but "do these correct functions actually work when connected."

The whole thing, as a user

An end-to-end test drives the whole system the way a real user would — clicking through the actual interface, submitting the form, waiting for the page that comes back — with every real part running behind it. It exercises the full stack at once, so passing one gives the strongest confidence available that the software actually works for the person using it.

That realism is expensive. An end-to-end test is the slowest kind, because the entire system has to be running and responding, and it is the most brittle: it fails for reasons that have nothing to do with a real defect — a slow response, a moved button, a timing hiccup. Tests that cry wolf get ignored, which is the worst thing a test can do.

Because they are slow and flaky, you keep the number of end-to-end tests to a bare minimum — enough to confirm the critical paths a user actually takes work start to finish, and no more. They sit at the top of the effort, checking the most and being trusted the least per test.

Many fast, few slow

the test pyramid: many at the base, few at the tip end-to-end tests integration tests unit tests the wide base fewer, slower, more realistic more, faster, cheaper
A three-band pyramid: a wide base of many fast, cheap unit tests, a narrower middle band of fewer integration tests, and a small tip of few slow, realistic end-to-end tests.

The test pyramid is a rule of thumb for how many of each kind of test to write. Picture the three kinds stacked by scope: a wide base of unit tests, a narrower band of integration tests in the middle, and a small tip of end-to-end tests on top. Write lots of the small fast ones, some of the middling ones, and very few of the broad slow ones. The idea comes from Mike Cohn, in his book on agile development.

The shape falls out of the trade the tests make. Lower in the pyramid, tests are small, isolated, and fast, and cheap enough to run constantly; higher up, each test drags in more of the system, so it grows slower, harder to write, and more brittle while it grows more realistic. You are buying realism with speed and reliability, and the pyramid says buy most of your coverage where it is cheapest.

Inverting it is the classic mistake — a pile of slow end-to-end tests on a thin base of unit tests. Such a suite is slow to run and quick to break, so it gets run rarely and trusted little, which is the opposite of what a suite is for. Wide at the bottom, narrow at the top, keeps the whole thing fast enough to actually use.

Arrange, act, assert

one test, three steps, then a verdict arrange set up inputs and state act call the code under test assert compare to expected then result matches no match pass result is as expected fail signals a defect
A test runs as three sequential steps: arrange sets up inputs and state, act calls the code under test, assert compares the result to the expected value, and the assertion then either passes when they match or fails when they do not, signalling a defect.

Almost every test, whatever its scope, has the same three parts. First you arrange: set up the situation the test needs — the inputs, the starting state, whatever the code expects to find. Then you act: call the code under test, the one function or path this test is about, and capture what it produces. Finally you assert: compare that result against the value you expected, and this last step is the whole point of the exercise. The three parts are why the pattern is remembered as arrange, act, assert.

Working through a test this way is like cooking a dish - you lay out the ingredients, cook it, then taste the result against how it was supposed to turn out. The assertion is the taste at the end: an assert states what the result should be, and if the actual result matches, the test passes silently. If it does not match, the assertion fails, and a failing assertion is exactly how a test reports that the code did something other than what it should have.

Keeping the three parts distinct is what makes a test readable. Anyone reading it can see the setup, the single thing being exercised, and the expectation being checked, without untangling which line does what. A test that mixes them together still runs, but it stops explaining itself, and a test you cannot read is one you will not trust when it fails.

Write the test first

red, green, refactor: a repeating cycle red write a failing test green minimal code to pass refactor clean up, test stays green write code until it passes now improve it next bit: a new failing test
The test-driven development cycle loops through three states: red, a newly written failing test; green, the minimal code that makes it pass; and refactor, cleaning the code while the test stays green, before looping back to a new failing test.

Test-driven development turns the usual order around: you write the test before the code it tests. First write a test for the next bit of functionality you want, and run it — it fails, because the functionality does not exist yet. Then write just enough code to make that test pass, nothing more. Then, with the test still passing, clean up what you wrote. This loop is remembered by the state of the tests at each step: red, green, refactor.

Writing the failing test first is like agreeing what finished looks like before you start building, so you know the exact moment you have reached it: it forces you to say exactly what "working" means before you build anything, so you know the precise moment you have reached it — the red test turning green. Minimal code to pass keeps you from building things no test asked for, and running the test first proves it can actually fail, so a later pass means something.

The refactor step is the one people skip, and skipping it is the most common way to get TDD wrong. Once the test is green you have a safety line: you can restructure the code to make it clean and still know instantly if you broke the behavior, because the test would go red again. Drop the refactor and you accumulate passing-but-messy code; keep it, and every cycle leaves the code both working and tidy.

Changing code without fear

Testing is load-bearing because it is what lets a project change without fear. Software that people use is edited constantly — bug fixes, new features, cleanups — and every edit risks breaking something that already worked. A trustworthy suite of tests turns each edit from a gamble into a checkable claim: make the change, run the tests, and let the green result stand in for the confidence you cannot get by reading the code and hoping.

That confidence compounds when more than one person is involved. On a shared project, you edit code you did not write and did not fully understand, and so does everyone else. The tests are the shared contract that says what the code is supposed to do, catching the moment someone's change contradicts it — which is why tested code is far safer to hand to a collaborator, or to come back to yourself six months later.

The same tests do double duty once a project matures: automation runs them on every change, so nobody has to remember to. That is a subject of its own, but the groundwork is the same either way — a test is only useful if it is honest and it is actually run, and the work is writing tests worth running.