Prefer to click through the interactive model?

Partial derivatives — study guide

The same fragments the interactive model serves, read in order. One source, two views.

One input at a time

A derivative asks a single question of a single-variable function: how fast does the output change as the one input changes. That question stops being well-posed the moment a function takes more than one input. A room's temperature depends on both position and time; a company's profit depends on both price and quantity sold. Ask "how fast is it changing" about a function like that and the honest answer is "changing with respect to what — and what are you holding still while it changes?"

A partial derivative answers exactly that narrower, answerable question: freeze every input except one, and ask the ordinary single-variable rate of change question about the input left moving. partial does not mean incomplete. It means restricted to one direction of change at a time, with every other direction pinned in place so the question has a single, unambiguous answer.

Everything here builds directly on the single-variable derivative already in hand: the same limit, the same rules for differentiating a term, applied one input at a time to a function of several.

Notation and the formal definition

The standard notation for a partial derivative is ∂f/∂x, read "partial f partial x." The rounded ∂ deliberately looks like, but is not, the ordinary d from single-variable calculus — the different symbol is a signal that something is being held fixed. The same quantity also shows up written f_x or D_x f; all three notations name the same object and are used interchangeably across textbooks.

The formal definition mirrors the single-variable derivative exactly, with one addition: everything except the input being varied is frozen inside the limit. For f(x, y), the partial derivative with respect to x at a point (a, b) is ∂f/∂x(a, b) = lim(h→0) [f(a + h, b) − f(a, b)] / h — the value of b never moves as h shrinks toward 0. Swapping the roles gives ∂f/∂y: hold x fixed at a and let y do the moving instead. Each partial derivative is genuinely just an ordinary single-variable limit, taken on the function you get by nailing down every input but one: like a control panel of several independent dials, where turning exactly one dial and watching a single readout move tells you that dial's effect, with every other dial pinned in place.

The ∂ symbol itself is younger than the ideas it now stands for. Nicolas de Condorcet introduced it in 1770 for a related partial-differential idea, Adrien-Marie Legendre adopted it for the partial derivative in 1786, and it was Carl Gustav Jacob Jacobi's use of it in 1841 that made the notation stick.

Computing a partial derivative

Computing a partial derivative needs no new differentiation rules. Treat every variable except the one you are differentiating with respect to as a plain constant, then apply the same power, sum, and product rules already familiar from a single-variable derivative. The only new step is deciding, term by term, which symbols are the constants for this particular pass.

Take f(x, y) = 3x^2*y − 4y^3 + 5. To find ∂f/∂x, treat y as a fixed number: 3x^2*y differentiates like 3y times x^2, giving 6x*y; −4y^3 is entirely constant with respect to x, so it differentiates to 0; the lone 5 also differentiates to 0. That leaves ∂f/∂x = 6x*y. To find ∂f/∂y instead, treat x as the fixed number: 3x^2*y is 3x^2 times y, so it differentiates to 3x^2; −4y^3 differentiates the ordinary way to −12y^2; the 5 again contributes 0. That leaves ∂f/∂y = 3x^2 − 12y^2.

At the point (2, 1), both partials are exact, ordinary numbers: ∂f/∂x(2, 1) = 6(2)(1) = 12 and ∂f/∂y(2, 1) = 3(2^2) − 12(1^2) = 12 − 12 = 0. The function itself is f(2, 1) = 3(2^2)(1) − 4(1^3) + 5 = 13. Two different numbers, 12 and 0, both describing the same point — because they answer two different questions about it.

What the slope means on a surface

contour lines of z = f(x, y), seen from above z=1 z=2 z=3 z=4 z=5 y held fixed at y0 — x increasing wide gap gently sloped here narrow gap steep here
Holding y fixed and cutting straight across the contour lines of z = f(x, y): where the contour lines sit close together the cut climbs quickly and the partial derivative with respect to x is large, where they sit far apart it climbs slowly.

A function of two inputs, z = f(x, y), can be pictured as a surface: every (x, y) pair sits on the ground, and z is the height above it. Fixing y at some constant value b and letting x alone move slices that surface along a single flat plane, and what is left on that cut is an ordinary single-variable curve — like cutting one flat slice through a loaf at a fixed position along its length, exposing a single two-dimensional cross-section curve on that cut face. That curve is g(x) = f(x, b), a function of x alone, exactly the kind of curve the ordinary derivative already knows how to handle.

∂f/∂x at a point is nothing more exotic than the tangent slope of that sliced curve at that point. Fixing x instead and letting y move slices the surface along the other direction, and ∂f/∂y is the tangent slope of that second curve. Two different slices through the same point on the same surface, two different slopes, both honest descriptions of how steep the surface is — just in different directions.

The figure shows the same idea from directly above, as a contour map: each line traces the points where the surface sits at one fixed height. Holding y fixed and moving x traces a straight cut across those lines. Where the lines sit close together the surface climbs fast over a short horizontal distance, so ∂f/∂x is large there; where they spread out, the same horizontal step buys far less height, and ∂f/∂x is small.

For f(x, y) = 3x^2*y − 4y^3 + 5 at the point (2, 1), the two slices disagree sharply: holding y at 1 and moving x, the surface climbs at slope ∂f/∂x(2, 1) = 12; holding x at 2 and moving y instead, the surface is momentarily flat, ∂f/∂y(2, 1) = 0. Both are correct descriptions of the exact same point.

Second partials and the order question

two orders of stepping from one point to another (x0, y0) (x0+h, y0) (x0, y0+k) (x0+h, y0+k) step x (∂/∂x) step y (∂/∂y) step y (∂/∂y) step x (∂/∂x) x then y (∂²f/∂y∂x), or y then x (∂²f/∂x∂y) — one destination
Stepping in x then y, or in y then x, both reach the same corner point — the two orders of a mixed second partial derivative that agree there whenever both mixed partials are continuous.

A partial derivative can itself be differentiated again, and now there is a choice to make about which input moves next. Differentiating ∂f/∂x with respect to x again gives the second partial ∂²f/∂x²; differentiating it with respect to y instead gives the mixed partial ∂²f/∂y∂x — step in x first, then y (the operator written nearest f acts first). Starting from ∂f/∂y and stepping in x gives ∂²f/∂x∂y — the same two steps, taken in the opposite order.

Those two mixed partials, ∂²f/∂x∂y and ∂²f/∂y∂x, are not automatically the same number; they are computed by different routes and only provably agree under a stated condition. Clairaut's theorem gives that condition: when both mixed partials exist and are continuous at and around a point, the order of differentiation does not matter there, and the two routes land on the same value. The theorem carries Alexis Clairaut's name from the eighteenth century, though the first fully rigorous proof came later, from Hermann Schwarz in 1873.

For f(x, y) = x^2*y^3, both orders agree everywhere. Starting from ∂f/∂x = 2x*y^3 and differentiating with respect to y gives ∂²f/∂y∂x = 6x*y^2. Starting instead from ∂f/∂y = 3x^2*y^2 and differentiating with respect to x gives ∂²f/∂x∂y = 6x*y^2 — the identical expression. That agreement is guaranteed here because f is a polynomial, and a polynomial's partial derivatives of every order are continuous at every point.

Where this goes wrong

The most common slip is differentiating the variable that was supposed to stay fixed. Given f(x, y) = 3x^2*y, finding ∂f/∂x means treating y as a plain number, so the whole term differentiates to 6x*y. Treating y as if it were also changing with x — and reaching for a rule like the product rule to differentiate x^2 and y together — answers a different question than the one that was asked, and gets a wrong term-by-term result even though every individual differentiation step looks legitimate.

A related but sharper mistake is erasing the fixed variable before differentiating, rather than treating it as a constant. In f(x, y) = 3x^2*y − 4y^3 + 5, the term 3x^2*y still has an x in it and still needs differentiating with respect to x; only −4y^3 and 5, which contain no x at all, actually vanish. Mentally deleting every symbol other than x first, instead of deleting only the terms that end up with a zero derivative, drops real terms that were still supposed to contribute.

A third mistake shows up once the two inputs are no longer independent — when y itself depends on x along some path, rather than sitting fixed. A partial derivative always answers the narrower question of one input moving while the others hold still; it does not by itself describe how f changes as several inputs move together along a path. Treating a partial derivative as if it already accounted for every input's motion is a category error, not an arithmetic one, and it is worth naming even though combining several partials into that broader answer is a separate idea.

Why every multivariable model needs this

Any system with more than one adjustable input needs a rate of change for each input separately before it needs anything more elaborate. A driving controller with a separate setting for steering and for speed changes differently depending on which setting moves; a multi-step pipeline with several independent tunable parameters behaves the same way. Asking "how sensitive is the output" without saying which input is moving is exactly as underspecified for a system like that as it was for a bare function of x and y.

That is the whole job of a partial derivative: one clean, answerable number per input, with everything else held still while it is measured. Nothing about combining several inputs' partials into a single, larger description of the system's overall sensitivity is decided here — that is real machinery built on top of exactly this, one input's rate of change at a time.

What carries forward is the discipline, not a formula: before asking how a many-input system responds, first pin down what every input except one is doing. Every partial derivative computed the way this concept describes is already a usable, complete answer to a real, narrower question, whether or not it is ever combined with anything else.