Prefer to click through the interactive model?

Optimization — study guide

The same fragments the interactive model serves, read in order. One source, two views.

Where the slope flattens

A function's value climbs, falls, climbs again — and at every smooth peak and every smooth valley, one thing is true just before the turn: the tangent line is flat. Optimization is the discipline of using that single fact backward, hunting down the input where a curve stops rising or stops falling, because that input is where the function does its best or its worst.

The tool doing the hunting is the derivative, which already reports a function's slope at any point. Set that slope to zero and you have named every place a peak or a valley COULD sit — not a full answer yet, only a short list of suspects. The rest of this concept is the discipline of separating the suspects that really are extrema, a rise-then-fall or a fall-then-rise, from the ones that only look like they should be, and of handling functions too tangled for that shortlist to work at all.

None of this is exotic. Every guided build ahead that tunes a physical system until it performs its best, sizes a pipeline so it wastes the least, or steers a small vehicle toward the fastest lap is, underneath, a version of the same question: given a way to measure "better", where is the input that makes it best? Finding maxima and minima by following the slope to where it flattens is the entire idea; everything else here is precision about how far to trust it.

Critical points, candidates not guarantees

f(x) = x^2 f'(0)=0 — minimum f(x) = -x^2 + 2.25 f'(0)=0 — maximum f(x) = x^3 / 2 f'(0)=0 — not an extremum
Three curves each with a horizontal tangent line where the derivative is zero: the first curve turns upward into a minimum there, the second turns downward into a maximum there, and the third has the same flat tangent but keeps rising straight through the point, showing that a zero derivative marks a candidate extremum, not a guaranteed one.

Call a point a critical point of a function if the derivative there is zero, or if the derivative fails to exist at all. That is the whole definition — it says nothing yet about whether the point is a peak, a valley, or neither. A critical point is a candidate for an extremum, nominated by the algebra, not a confirmed one.

The figure shows why the distinction matters. Take f(x) = x^2: its derivative is 2x, zero only at x = 0, and the curve genuinely turns there, bottoming out into a minimum. Flip the sign to f(x) = -x^2 and the same zero-derivative point becomes a maximum instead, the curve turning the other way. Both are honest extrema, found by the same test.

Now take f(x) = x^3. Its derivative is 3x^2, also zero at x = 0 — a critical point by the same rule. But 3x^2 never goes negative; it sits at zero for an instant and climbs right back to positive on both sides. The curve does not turn at all, it just pauses its climb and keeps climbing. x = 0 is a critical point of x^3 that is not a maximum, not a minimum, nothing — the algebra nominated a candidate the curve rejected. Every method in this concept for actually judging a critical point exists because that rejection is possible.

The first-derivative test

f(x) = x^3 - 3x — sign of f' on either side of a critical point f(x) x x = -1 x = 1 local max local min f' positive f' negative f' positive
The curve f(x) = x^3 - 3x with the sign of its derivative marked on either side of each critical point: positive then negative around x = -1, marking a local maximum, and negative then positive around x = 1, marking a local minimum.

A critical point earns extremum status by what the slope does around it, not at it. The first-derivative test reads the sign of the derivative just to the left and just to the right of a critical point: rising into it and falling away, positive then negative, makes it a maximum; falling into it and rising away, negative then positive, makes it a minimum; the same sign on both sides means the curve never actually turned, so the point is neither.

Run the test on f(x) = x^3 - 3x, whose derivative f'(x) = 3x^2 - 3 is zero at x = -1 and x = 1. Just left of x = -1 the derivative is positive, the curve is climbing; just right of it, the derivative is negative, the curve is falling. Positive to negative means x = -1 is a local maximum, and the curve's value there, f(-1) = 2, is that peak's height. Cross x = 1 the opposite way, negative on the left and positive on the right, and the sign change identifies a local minimum, with f(1) = -2 as the valley's depth.

The test only needs the sign, never the exact slope value, which is what makes it work even where the second-derivative test, checking curvature instead of surrounding sign, goes silent.

The second-derivative test

f(x) = x^3 - 3x — concavity either side of each critical point f(x) x max, f''(-1)=-6 min, f''(1)=6 x=0, f''=0 — inconclusive concave down concave up
The curve f(x) = x^3 - 3x split by concavity: bending downward through the local maximum at x = -1, bending upward through the local minimum at x = 1, and passing through an inflection point at x = 0 where the second derivative is zero and concavity itself gives no information.

Concavity gives a shortcut to the same verdict the first-derivative test reaches by checking signs on either side. The second derivative measures how the slope itself is changing: where it is negative, the curve bends downward like the top of a hill, and any critical point sitting there is a local maximum; where it is positive, the curve bends upward like the bottom of a bowl, and the critical point is a local minimum.

On f(x) = x^3 - 3x, the second derivative is f''(x) = 6x. At the critical point x = -1, f''(-1) = -6, negative, the curve is bending downward there, so x = -1 is a local maximum. At x = 1, f''(1) = 6, positive, the curve is bending upward, so x = 1 is a local minimum. Curvature and the sign of the surrounding slope are two different measurements of the same curve, and here they agree.

The test has a blind spot: when the second derivative is also zero at a critical point, it reports nothing, not "neither an extremum", simply no information either way. f(x) = x^3 lands exactly there. Its second derivative f''(x) = 6x is zero at the same x = 0 where the first derivative is zero, so the second-derivative test is inconclusive at that point, even though x = 0 is, in fact, not an extremum. Settling a case like that falls back to checking the sign of the slope on either side directly.

Local best versus global best

g(x) = x^3 - 3x^2 — two local extrema, no global ones g(x) x to -infinity to +infinity x=0, local max (not global) x=2, local min (not global)
The curve g(x) = x^3 - 3x^2, with a local maximum at x = 0 and a local minimum at x = 2 marked, while the curve itself keeps falling toward negative infinity to the left and keeps rising toward positive infinity to the right, so neither marked point is a global extremum.

A local extremum only promises to be the best point in its immediate neighborhood, the highest or lowest value nearby, with no claim about anywhere else on the curve. A global extremum promises more: the single highest or single lowest value the function reaches, full stop, across whatever domain is under discussion. Every global extremum is automatically a local one, but plenty of local extrema are not global, and mixing the two up is a real and common mistake, like a hiker who reaches the top of the nearest hill and, standing right there, has no way to tell whether a taller peak sits somewhere else in the range.

Take g(x) = x^3 - 3x^2. Its critical points sit at x = 0 and x = 2, with g(0) = 0 a local maximum and g(2) = -4 a local minimum, each confirmed the same way any critical point is: a sign change in the first derivative, matched by the curvature the second derivative reports. But follow g further out and neither holds up as global. As x grows without bound, g(x) grows without bound too, climbing past 0 and never turning back, so no point can be the global maximum. As x falls without bound, g(x) falls without bound as well, sinking past -4 forever, so no point can be the global minimum either.

So this curve has two real local extrema and zero global ones — a function is not required to have a best or worst value at all once its domain runs open in both directions. Whether a global extremum exists, and where, depends on the domain as much as on the function; restrict g to a closed stretch of the line instead of the whole number line, and the picture changes completely.

Closed intervals, checking the edges too

g(x) = x^3 - 3x^2 on [-1, 4] — every candidate, endpoints included g(x) x x=-1, g=-4 global min (tied) x=0, g=0 x=2, g=-4 global min (tied) x=4, g=16 global max
The curve g(x) = x^3 - 3x^2 restricted to the closed interval from -1 to 4, with all four candidates marked: the endpoints and the two interior critical points. The global maximum sits at the right endpoint x = 4, and the global minimum is tied between the left endpoint x = -1 and the interior critical point x = 2.

Restrict a continuous function to a closed, bounded interval, a stretch of the line with both ends included, and the picture that seemed hopeless turns solid: the function is now guaranteed to attain both a global maximum and a global minimum somewhere on that interval. That guarantee is the Extreme Value Theorem, and it costs nothing beyond continuity and a fenced-in domain. The catch is that "somewhere" might not be an interior critical point at all — it might be sitting right at one of the two ends, like measuring a fenced field for its highest and lowest ground: you have to walk the interior and check right along the fence line, because the actual high or low point might be sitting on the boundary.

The closed-interval method turns that guarantee into a procedure: list every interior critical point, add both endpoints, evaluate the function at all of them, and the largest value is the global maximum while the smallest is the global minimum. Take g(x) = x^3 - 3x^2 again, this time boxed into the interval from -1 to 4. The interior critical points are still x = 0 and x = 2, giving g(0) = 0 and g(2) = -4. The two endpoints add g(-1) = -4 and g(4) = 16. Comparing all four: the global maximum on this interval is 16, at the endpoint x = 4, nowhere near either interior critical point. The global minimum is -4, and it is tied, reached at both the endpoint x = -1 and the interior critical point x = 2.

Drop either the closed-ness or the boundedness of the interval and the guarantee disappears with it: an open interval can let the function creep toward a value it never actually reaches, and an unbounded one can let it run off forever, exactly as it did without any interval at all.

Flat in every direction

z = x^2 + y^2 — a bowl gradient zero at center — the minimum z = x^2 - y^2 — a saddle gradient zero at center — neither max nor min
Two contour diagrams sharing a center point of zero gradient: closed rings around a bowl-shaped surface, where the center is genuinely the lowest point, beside crossing hyperbola-shaped level curves around a saddle-shaped surface, where the center is flat in every direction but is neither a maximum nor a minimum.

A function of two variables has a slope in every direction at once, and the gradient collects all of them into a single vector that points the way uphill. A critical point of such a function is any input where that gradient is the zero vector, flat in every direction simultaneously, not just along one axis. It is the same idea as a one-variable critical point, promoted to more dimensions: the algebra nominates a candidate, and nothing about the nomination guarantees what kind of candidate it is.

For a function shaped like a bowl, z = x^2 + y^2, the gradient is zero only at the origin, and the origin is genuinely the lowest point on the whole surface, every direction you could step leads upward. But a critical point can also be a saddle point: flat in every direction at that one spot, yet rising along one direction and falling along another, so it is neither a local maximum nor a local minimum of the surface. z = x^2 - y^2 has exactly this shape at the origin: walk along the x-axis and the surface climbs, walk along the y-axis and it drops, and the gradient at the center is zero either way.

The lesson carries over unchanged from a single variable: a zero gradient marks a critical point, and a critical point is still only a candidate. Confirming what kind it is takes more machinery than checking that the gradient vanished, the same caution a curve as simple as x^3 already demanded in one dimension.

When there's no clean formula, stepping downhill

Setting a derivative to zero and solving only works when the algebra cooperates — plenty of real functions have derivatives too tangled to solve for an exact critical point by hand, or no closed-form solution at all. For those, optimization gives up on finding the answer in one step and instead walks toward it, like a ball released partway up a valley wall: it rolls in the direction the ground slopes down, picks up speed on a steep stretch, and comes to rest wherever the ground finally goes flat.

The idea: start somewhere, measure the slope right there, and take a small step in the direction the slope goes down, since a positive slope means the function falls to the left and a negative slope means it falls to the right. Measure the slope again at the new spot and step again. Each step is scaled by a fixed step size, small enough to avoid overshooting, large enough to make real progress, and the process repeats until the slope goes flat enough that another step would barely move the value at all. That stopping condition is the same "flat means done" idea running through this whole concept, just found by walking instead of by algebra.

The step size controls the whole behavior. Too small and the search crawls, needing far more steps than the problem deserves. Too large and a step can fly straight past the bottom, land somewhere with an even steeper slope than before, and start overshooting worse each time, so the search diverges instead of settling. And the search only ever "sees" the slope right where it is standing, so on a curve with more than one dip, a shallow one can trap it well short of the true lowest point elsewhere, the same local-versus-global distinction, now showing up as a failure mode instead of an abstract case. The model on this page runs exactly this search: set a starting point and a step size and watch it settle, overshoot, or get stuck.

Why this matters ahead

Every guided build ahead that talks about tuning, sizing, or steering toward the best result is running some version of the reasoning in this concept, even when it never says the word optimization out loud. Point a system at a number worth minimizing or maximizing, lap time, wasted capacity, error, and the question is always the same one asked here: where does the slope of that number, with respect to whatever you can adjust, go flat?

Closed-form critical points handle the tidy cases, where the relationship between an adjustable input and the outcome is simple enough to write down and differentiate directly. Most real systems are not that tidy, which is exactly why a stepping approach, measure the slope where you stand, take a step against it, repeat, turns out to matter far beyond any one worked example. Run that same idea over many adjustable parameters at once, instead of just one, and it is usually called gradient descent; a later concept picks that name up and runs with it in a very specific setting.

For now, the two habits worth keeping are the same ones this concept keeps returning to: treat a flat slope as a candidate, never a guarantee, and check the edges of whatever range you are actually allowed to search before declaring a winner.