Prefer to click through the interactive model?

Gradients — study guide

The same fragments the interactive model serves, read in order. One source, two views.

One vector, all the partials

Take a function of more than one input and ask for its partial derivatives, one with respect to each input in turn, and several numbers come back separately: a rate for x, a rate for y, and so on for every input the function takes. Line those numbers up in a fixed order and they form a single object built out of all of them at once. That object is the gradient, a vector with one component per input.

A gradient reports a direction and a size in the same breath. Stand at any point where the gradient exists and is not the zero vector, and it points the way to step for the function to grow fastest, and its length says exactly how fast the function grows per step in that direction. Neither claim is approximate or a rule of thumb; both follow directly from the same partial derivatives that assembled the vector in the first place.

The rest of this concept works through what that direction and that size actually mean, how to build the vector from a function stated in ordinary algebra, and why a right angle keeps turning up wherever a gradient does.

Building the vector from the partials

the gradient as a sum of two components x y df/dx = 4 df/dy = 2 grad f = (4, 2)
The gradient vector (4, 2) at a point on f(x,y) = x^2+y^2 is the vector sum of its two partial-derivative components, 4 along x and 2 along y, completed as a rectangle.

Building a gradient is mostly bookkeeping once the partial derivatives are already in hand. Take a function of two inputs, x and y, and compute its two partial derivatives separately, holding the other input fixed each time exactly as ordinary partial-derivative work already does. Write the x-partial first, then the y-partial second, and that pair, in that order, is the gradient: (df/dx, df/dy).

Nothing about the order is arbitrary once a function is fixed: the first slot always answers how fast the function changes per unit of x, the second always answers the same question for y, and swapping them would silently swap which axis the resulting vector points along. For a function of three or more inputs the pattern repeats without change, one partial derivative per input, listed in the same fixed order every time.

The figure on this page assembles one concrete gradient this way. For the function f(x, y) = x^2 + y^2 evaluated at the point (2, 1), the x-partial works out to 4 and the y-partial to 2. Laid out as two component vectors and completed into a rectangle, those two components sum to the single vector (4, 2), the gradient at that point, formed by nothing more than an ordinary vector sum.

The direction of fastest increase

gradient crosses to the next level curve f = 16 f = 9 (3, 0) fastest increase zero rate of change
At the point (3, 0) on the circular level curve where f(x,y) = x^2+y^2 equals 9, the gradient vector points radially outward to the next level curve, f = 16, along the shortest path, while the tangent direction along the same level curve has zero rate of change.

Stand at a point where a function is defined and imagine every possible direction to step in next, not just along the x-axis or the y-axis but any direction at all. Each direction has its own rate of change, called a directional derivative, and every one of those rates can be read straight off the gradient: multiply the gradient by the unit vector for that direction and add the results, the same multiply-and-add recipe any two vectors combine by.

Among every direction available, exactly one produces the largest directional derivative, and it is the direction the gradient itself points in — like a hiker on a hillside turning to face the exact direction that climbs the mountain fastest from where they stand. Turn to face any other direction and the rate of increase only gets smaller; turn to face squarely opposite the gradient and that same rate becomes the steepest possible rate of decrease instead.

That claim holds at every point where the gradient is not the zero vector: the gradient is not merely correlated with the direction of fastest increase, it is that direction, exactly, not approximately. A function's steepest uphill direction and its gradient's direction are the same fact stated two ways.

How fast is fastest

Direction is only half of what a gradient reports. The other half is a plain number: the gradient's length, its magnitude, computed the same way any vector's length is, by squaring each component, adding the squares, and taking the square root. That number is not a side detail; it is the actual value of the fastest possible rate of increase, the largest directional derivative any direction could produce at that point.

A short gradient means the function is nearly flat in every direction near that point. Step wherever you like and the function barely moves, because even the best direction available only manages a small rate of change. A long gradient means the opposite: the function is climbing steeply in its best direction, and stepping in any other direction still catches a good share of that steepness, just less of it.

Zero is the extreme case. A gradient of length zero has no direction to report at all, because — for a smoothly varying function — every directional derivative through that point works out to zero, and the function is momentarily flat in every direction simultaneously, the same way a flat spot on a curve of one variable generalizes once a second input enters the picture.

Crossing the contour lines squarely

gradient meets tangent at a right angle level curve: x^2 + y^2 = 9 (3, 0) grad f = (6,0) tangent direction
At the point (3, 0) on the level curve x^2 + y^2 = 9, the gradient vector (6, 0) and the tangent direction to the level curve meet at an exact right angle.

A level curve collects every point where a function of two variables holds one fixed value, every (x, y) with f(x, y) = 9, say, forming a curve of its own. Walk along a level curve and the function's value never changes at all, by definition; the rate of change in that tangential direction is exactly zero, no matter which level curve or which function is involved.

At any point where the function is differentiable and its gradient is not the zero vector, the gradient is perpendicular to the level curve passing through that point — like walking straight across a topographic map's contour lines instead of following one, the fastest way to change elevation. The two facts fit together rather than sitting as two separate rules: the tangential direction along the level curve is the one direction with zero rate of change, the gradient's direction is the one direction with maximum rate of change, and a maximum and a zero pulling as far apart as two directions in a plane can is exactly what perpendicular means.

The condition matters as much as the claim. A flat point where the gradient is the zero vector has no single perpendicular direction to point in, and a function that fails to be differentiable at a point can break the perpendicularity outright. Away from those exceptions, though, the right angle is not a coincidence belonging to any one function; it falls straight out of the fastest-increase property already established.

Computing one, exactly

Take the function f(x, y) = x^2 + y^2 and find its gradient at the point (3, 4). The two partial derivatives are df/dx = 2x and df/dy = 2y, worked out once for the function in general and then evaluated at the point in question: 2*3 = 6 and 2*4 = 8. Stacked in order, the gradient at (3, 4) is the vector (6, 8).

Its magnitude follows the same square-add-root recipe as any vector's length: 6^2 + 8^2 = 36 + 64 = 100, and the square root of 100 is 10. That 10 is the exact maximum rate of increase available at (3, 4); no other direction from that point produces a directional derivative larger than 10.

The perpendicularity claim checks out just as concretely at a simpler point on the same function. At (3, 0), the gradient is (6, 0), pointing straight along the x-axis, and the level curve through that point is the circle x^2 + y^2 = 9. The tangent line to that circle at (3, 0) runs straight up and down, parallel to the y-axis, meeting the horizontal gradient at a right angle with nothing approximate about it.

Where this vector shows up next

A gradient answers a question that comes up anywhere a quantity depends on more than one adjustable input: which way, and how strongly, does that quantity change fastest right here. A system steering toward a target, a process tuning several settings at once, or a search hunting for the best combination of inputs to some measurable outcome all eventually need that same answer, computed at whatever point they currently sit.

None of those systems need anything more exotic than what this concept already covers to ask the question. What they do with the answer, whether they step with the gradient, against it, or use it to decide when to stop adjusting anything at all, is a separate concern layered on top of the vector itself.

The vector does not change meaning depending on who asks. A machine deciding which way to nudge its next move and a much larger search deciding which way to nudge a long list of settings are both reading the same object: partial derivatives, one per input, pointing the way some quantity climbs fastest from exactly where it currently stands.