Read this as a study guide instead

Debugging

Reproducing, isolating, and fixing: hypothesis testing on your own code.

Code with its results

A notebook-style walk through the idea — every output shown is the real result of the code above it.

Watch a bug, corner it, fix it

See a function give the wrong answer, add prints to inspect its state, read the exception it raises, and fix the off-by-one.

Real output Every Out block below was produced by running the code above it. You can copy these cells into your own python3 and run them top to bottom to see the same numbers.

In the tracker, a bug is code doing what you wrote, not what you meant. Watch a rule meant to total a ticket's hours return the wrong number.

In [1]
def total_hours(tickets):
    s = 0
    for i in range(len(tickets) - 1):
        s = s + tickets[i]
    return s

print(total_hours([4, 6, 3]))
Out [1]
10

The total is wrong but nothing crashed. Add a labeled print inside the loop to watch which tickets it actually counts.

In [2]
def total_hours(tickets):
    s = 0
    for i in range(len(tickets) - 1):
        print("counting ticket", i, "hours", tickets[i])
        s = s + tickets[i]
    return s

print(total_hours([4, 6, 3]))
Out [2]
counting ticket 0 hours 4
counting ticket 1 hours 6
10

The print shows the last ticket, index 2, is never counted: range(len(tickets) - 1) stops one short. A different mistake crashes, so reach past the queue and catch it.

In [3]
queue = ["BUG-12", "BUG-15", "BUG-19"]
try:
    print(queue[3])
except IndexError as e:
    print("IndexError -", e)
Out [3]
IndexError - list index out of range

IndexError - list index out of range is exactly what a traceback's last line reports: the exception type and a one-line why. The fix for the total is to iterate the hours, leaving no index to miscount.

In [4]
def total_hours(tickets):
    s = 0
    for hours in tickets:
        s = s + hours
    return s

print(total_hours([4, 6, 3]))
Out [4]
13

The same ideas, as prose

These are the exact fragments the model serves — also available as an ordered study guide.

A bug is a hypothesis you have not tested

reproduce hypothesize test one change fix then isolated not yet isolated: a new hypothesis
The debugging loop: reproduce, then hypothesize, then test one change; if the cause is not yet isolated the test feeds a new hypothesis, and once isolated you fix.

A bug is not a mystery; it is code doing exactly what you wrote instead of what you meant. Debugging is the work of finding the gap between the two, and it has a shape that never changes: reproduce the failure, isolate its cause, fix it. The staring-and-guessing that beginners do is not a faster version of this loop — it is skipping the loop, and it is why the same bug can eat an afternoon.

The reliable move is to treat every bug as a hypothesis you have not yet tested like a doctor with a symptom: you do not prescribe a cure on sight, you form a diagnosis and run the one test that would confirm or rule it out. You have a symptom — a wrong number, a crash, a blank screen — and a guess about the cause. Debugging is making one change that would confirm or refute that guess, watching what happens, and using the result to sharpen the next guess. Each pass narrows the search until the cause is cornered and there is only one thing left it could be.

None of this requires talent, and almost none of it requires cleverness. It requires being systematic when the instinct is to thrash. The rest of this concept is the loop in detail: how to make a failure repeat, how to read what the language already tells you, how to watch a program's real state, and the handful of bugs you will meet again and again.

Make it happen on demand

The first job is not to fix the bug; it is to make it happen whenever you want. A failure you cannot reproduce, you cannot investigate, because you have no way to tell whether a change helped, hurt, or did nothing. "It broke once yesterday" is a rumor. "It breaks every time I run this with an empty list" is a bug you can work on.

Reproducing means pinning down the exact inputs and steps that trigger the failure, then confirming it recurs when you repeat them. A bug that appears only sometimes is telling you something too: the missing ingredient is usually a piece of state you have not accounted for — an order of events, a leftover file, a value that is sometimes empty. Finding what makes it intermittent is often most of the fix.

Once it recurs, shrink it. Strip away everything that is not needed to still see the failure — fewer inputs, smaller data, less code around it. A bug that reproduces with a three-line function and one short list is far easier to reason about than the same bug buried in a full program. The smallest case that still fails is the cleanest signal you will get, and it is worth the few minutes it takes to carve out.

Read the traceback bottom-up

Traceback (most recent call last): File "report.py", line 9, in main show(scores) File "report.py", line 4, in show print(scores[3]) IndexError: list index out of range read bottom-up the call stack: frames, newest last read first: the exception type and message
A traceback is read bottom-up: the frames above are the call stack, newest last, and the final line names the exception type and its message.

When Python hits an error it cannot handle, it prints a traceback and stops. A traceback looks like a wall of text, and the instinct is to skim it or panic, but it is a precise report and it is read from the bottom. The very last line names the exception type and a short message — IndexError: list index out of range, TypeError: can only concatenate str (not "int") to str. That one line is the WHAT, and it is usually enough to know what kind of mistake you made.

Above that last line is the block that tells you WHERE. It lists the source locations the program passed through on its way to the error, each with a file, a line number, and the actual line of code, printed one under another. The bottom of that block is where the exception was raised; reading it first, before anything else, is the single habit that turns a traceback from noise into a map.

The message is written for a human and rewards being read literally. "List index out of range" means you asked a list for a position it does not have. "Can only concatenate str to str" means you tried to add a string and a number. The words are not decoration — they name the exact rule you broke, and the line number tells you where you broke it.

Every frame is a paused call

The stack of locations in a traceback is not a random list; it is the chain of function calls that were still running when the error struck. Each entry is a frame: one call that had started but had not yet returned. The top frame is where the program began the doomed path, and each frame below is a call made by the one above it, down to the bottom frame where the exception was actually raised. Python labels it most-recent-call-last so you know which end is which.

This matters because your code rarely fails alone. The bottom frame is often inside a library you did not write — a Python builtin, a package — and the raw error there can look alien. The frame that helps you is usually the lowest one that names a file you own. That is the last decision your code made before handing control to someone else's, and it is where your fix almost always goes.

So the reading order is: last line for the exception, then walk up the frames until you reach your own code. A traceback that dives ten frames deep into a library is not ten times harder; it is the same two questions — what broke, and which of my lines led there.

Print what you assume

Most bugs come from a gap between what you believe a value is and what it actually is, and the fastest way to close that gap is to look. A print dropped into the code shows you the real state at a real moment — no tools, no setup, works anywhere output appears. It is the humblest debugging technique and the one experienced programmers reach for first, because it answers the only question that matters: is the program holding what I think it is holding?

Print a value together with a label, not on its own. print("count", count) tells you both the number and which number it is; a bare print(count) leaves you counting anonymous lines of output. When a TypeError is in play and a value is not behaving, print its type — print(type(value)) — because a "3" that looks like a number until it refuses to add is the most common version of this trap.

Placement is a search. Put a print where you are sure the state is still correct, and another where you suspect it has gone wrong; the failure lives between them. Move the prints closer together across a few runs and you bisect the code down to the first line where reality and your mental model part ways. That line is the bug, and you found it by looking instead of guessing.

The usual suspects

A handful of bugs account for a large share of the time programmers lose, and knowing their faces means recognizing them in seconds instead of hours. The first is the off-by-one: a loop or an index whose boundary is wrong by exactly one. range(len(items) - 1) quietly skips the last element; items[len(items)] reaches one past the end and raises IndexError. When a result is short by one, or a loop touches almost-but-not-quite everything, suspect the boundary. Iterating the values directly with for item in items sidesteps the whole class, because there is no index left to miscount.

The second is type confusion. A value that looks right can be the wrong type — a "3" read from input or a file is a string, not the number it resembles, and "3" + 4 raises TypeError rather than giving 7. Numbers arriving as text must be converted with int(...) or float(...) before arithmetic. When an operation complains about types, print the type of each operand and the culprit is usually obvious.

The third is subtler and famous: the mutable default argument. A default parameter value in Python is evaluated once, when the def runs, not fresh on every call. So a parameter that defaults to an empty list reuses the same list across every call that omits it, and values pile up between calls as if the caller were haunted. The fix is to default to None and build a new list inside the body. It catches nearly everyone once, which is exactly why it is worth meeting on purpose.

When prints are not enough

Print statements are perfect until they are not. When a bug hides deep in a loop, or you would need a dozen prints to see the shape of the state, a debugger earns its place. A debugger pauses the program at a line you choose and hands you the controls: you can inspect any variable at that exact moment, run the next line and stop again, step into a call to see inside it, or let it continue. Instead of guessing what to print in advance, you stop the program and ask it questions.

In Python the entry point is one builtin. Writing breakpoint() on a line drops you into pdb, the standard debugger, when execution reaches it — added to the language in version 3.7 so it works everywhere without an import. From the pdb prompt, p prints an expression, n runs the next line, s steps into a function call, and c continues until the next breakpoint or the end. A few minutes with these four commands replaces a scattering of prints and a rerun for each one.

A debugger is interactive by nature: it stops and waits for you to type, which is exactly why it belongs at a real terminal and not inside code meant to run untended. The mental model is the payoff — a program is not a black box that either works or does not; it is a sequence of states you are allowed to freeze and examine, one line at a time.

A skill, not a talent

The programmers who look like they never get stuck are not seeing something you cannot; they are being disciplined where the instinct is to flail. Debugging is a skill that trains, and the training is a habit: reproduce before you theorize, read what the language already told you, look at real state instead of imagined state, and change one thing at a time.

That last rule carries most of the weight. When a bug is frustrating, the temptation is to change three things at once and rerun, hoping something sticks — and when it works, or fails, you have learned nothing, because you cannot tell which change mattered like a controlled experiment: change one variable at a time so the result tells you which change actually mattered. Holding a single hypothesis and making a single change keeps the signal clean: the result points at one cause, and you either confirm it or cross it off and move on.

Debugging your own code well is also the quiet foundation for everything harder — writing tests, reading someone else's program, tracing a failure across a whole system. All of it is the same loop, scaled up. Learn to corner a bug in a three-line function and you have learned the move you will use for the rest of it.