A saddle point is, in multivariable calculus, a stationary point of a real-valued function that is neither a local maximum nor a local minimum: arbitrarily nearby points have both larger and smaller function values. Its name comes from the resemblance of a typical graph to a riding saddle, which curves upward in one direction and downward in another. In mathematical optimization, the term also denotes a point satisfying simultaneous minimization and maximization conditions in separate groups of variables. These meanings are related, but not equivalent. (ocw.mit.edu)
Definition and elementary examples
For a differentiable function , with open, the usual calculus definition requires a critical point , meaning that its gradient satisfies . It is a saddle point if every neighborhood of contains points such that
Thus, a zero gradient identifies a candidate, not its classification. The requirement of nearby values on both sides distinguishes a saddle from an extremum. (ocw.mit.edu)
The standard example is the polynomial
At the origin its gradient vanishes. Along , the function is , with a minimum at zero; along , it is , with a maximum there. Direct substitution therefore demonstrates saddle behavior. Rotating the coordinate axes changes its appearance but not the existence of nearby values above and below zero. (ocw.mit.edu)
The Hessian and second-derivative test
For a function with continuous second partial derivatives, the Hessian matrix describes its second-order behavior. In two variables, at a critical point , define
This is the determinant of the Hessian. If , the point is a saddle. If , it is a strict local minimum when , and a strict local maximum when . When , the test is inconclusive. (web.mit.edu)
In higher dimensions, linear algebra provides the corresponding criterion through the Hessian’s eigenvalues and eigenvectors. Both positive and negative eigenvalues imply a saddle: different directions have opposite quadratic curvature. All positive eigenvalues imply a strict local minimum; all negative eigenvalues imply a strict local maximum. A singular Hessian requires further investigation unless mixed signs already establish saddle behavior. The determinant alone generally cannot classify higher-dimensional critical points. (ocw.mit.edu)
Degenerate and strict saddles
A saddle is nondegenerate if its Hessian is nonsingular. Degenerate saddles can also occur: for example, direct calculation for
gives a zero Hessian at the origin, while the coordinate axes still exhibit opposite signs. This illustrates why a failed second-derivative test does not rule out a saddle. Higher-order terms or direct comparisons of function values may determine the classification. (web.mit.edu)
In optimization research, a strict saddle commonly means a stationary point with at least one strictly negative Hessian eigenvalue. This definition emphasizes a direction of negative curvature rather than nonsingularity. Consequently, a strict saddle may have zero eigenvalues; under this convention, some local maxima also satisfy the strict-saddle condition. Terminology must therefore be interpreted within its mathematical context. (arxiv.org)
Minimax problems and game theory
For , a saddle pair satisfies
With the other variable fixed, minimizes and maximizes. These inequalities imply equality of the attained minimax and maximin values. They are global conditions, stronger than simply being a nonextremal stationary point. A convex function in that is concave in is particularly suitable: an interior stationary pair satisfies the saddle inequalities under these assumptions. (stanford.edu)
In game theory, saddle pairs describe Nash equilibria of two-player zero-sum games. For a payoff matrix whose row player maximizes and column player minimizes, a pure-strategy saddle entry is smallest in its row and largest in its column. Some matrices have no such entry. Allowing mixed strategies extends the strategy spaces to probability distributions; finite zero-sum games then possess a saddle equilibrium in expected payoff. (ocw.mit.edu)
In constrained convex optimization, saddle conditions for the optimization Lagrangian connect primal solutions, multipliers, and Lagrangian duality. Under appropriate hypotheses, they characterize primal-dual optimality and relate to the Karush–Kuhn–Tucker conditions. (stanford.edu)
Numerical optimization
Saddles matter when minimizing an objective function, including a loss function in machine learning. An exact stationary point produces no update under ordinary gradient descent, while small gradients nearby can make progress slow. Negative curvature nevertheless identifies directions in which the objective decreases. (arxiv.org)
Avoidance results require explicit assumptions. For twice continuously differentiable objectives with Lipschitz-continuous gradients and sufficiently small fixed step sizes, gradient descent with an absolutely continuous random initialization has probability zero of converging to a specified strict saddle. This is not an unconditional guarantee of convergence, nor a bound on escape time. (arxiv.org)
Saddle-point approximation
In complex analysis, saddle points also organize approximations to integrals such as
Stationary points satisfy . The method of steepest descents deforms the contour, when analyticity and singularities permit, through relevant saddles along paths of constant imaginary phase. Local expansions around these points produce an asymptotic expansion for large . Which saddles contribute depends on the contour and parameters, not merely on the stationary-point equation. (dlmf.nist.gov)