The Hessian matrix is a square matrix containing the second-order partial derivatives of a scalar-valued function of several variables. It describes how the function’s first derivatives change and provides a local, second-order description of curvature. Named after the German mathematician Ludwig Hesse, it is a fundamental tool in multivariable calculus and mathematical optimization, especially for classifying stationary points and constructing curvature-based numerical methods. (web.mit.edu)
Definition and symmetry
Let , where is an open set in Euclidean space. At a point where the relevant second partial derivatives exist, its Hessian is written , , or . Using the convention
it is the Jacobian matrix of the gradient, regarded as a column vector. For two variables,
Diagonal entries measure changes in a coordinate’s own first derivative; off-diagonal entries measure changes involving two different coordinates. (docs.jax.dev)
If has continuous second partial derivatives in a neighborhood of the point, the mixed derivatives agree, so , where denotes matrix transpose. This regularity condition is sufficient for symmetry; merely having all second partial derivatives at a point does not guarantee it. In one variable, the Hessian reduces to the matrix containing . (web.mit.edu)
Local curvature and quadratic approximation
For a twice continuously differentiable function, the second-order expansion associated with the Taylor series is
The final term becomes negligible relative to as . Thus, the gradient supplies the linear part of the local approximation, while the Hessian supplies its quadratic form. For a fixed direction ,
This identity expresses the second directional derivative in terms of the Hessian. (math.ucr.edu)
For a symmetric Hessian, the spectral theorem gives orthogonal principal directions. Its eigenvalues and eigenvectors describe the signs and magnitudes of directional second derivatives along those directions. Positive eigenvalues indicate upward bending, negative eigenvalues downward bending, and zero eigenvalues directions whose behavior is not determined by the quadratic term alone. (math.ucr.edu)
Classification of stationary points
At an interior critical point satisfying , the Hessian determines the standard second-derivative test, assuming is twice continuously differentiable nearby:
- If is positive definite, meaning for every nonzero , then is a strict local minimum.
- If it is negative definite, then is a strict local maximum.
- If it is indefinite, with both positive and negative eigenvalues, then is a saddle point.
- If it is positive or negative semidefinite but not definite, the test is inconclusive.
A positive-semidefinite Hessian is necessary at a twice-differentiable interior local minimum, but it is not sufficient by itself. (math.ucr.edu)
In two variables, let
at the stationary point. If , the sign of distinguishes a minimum from a maximum; if , the point is a saddle; if , further analysis is needed. This determinant test does not extend unchanged to higher dimensions: the full definiteness of the Hessian matters. (mit.edu)
Convexity and examples
On an open convex set, a twice continuously differentiable function is a convex function if and only if its Hessian is positive semidefinite everywhere. A positive-definite Hessian everywhere implies strict convexity, but the converse fails: is strictly convex although . These conditions are central to convex optimization. (stanford.edu)
For the quadratic polynomial
direct differentiation gives and . Consequently, its Hessian is constant, and the quadratic approximation is exact. As concrete examples, has Hessian and a strict minimum at the origin, whereas has Hessian and a saddle there. (stanford.edu)
Numerical optimization and computation
In optimization, Newton’s method uses the Hessian to construct a quadratic model of an objective function. Its step solves the linear system
When the Hessian is positive definite, this step minimizes the local quadratic model. Unlike gradient descent, it accounts for curvature in different directions. A line search can control the step length because an accurate local model need not remain accurate far from the current point. (courses.csail.mit.edu)
A dense Hessian contains entries, making explicit storage expensive in large machine-learning problems, including curvature analysis of a loss function. Automatic differentiation can instead compute Hessian–vector products without constructing the whole matrix. For fixed and continuous second derivatives,
Alternatively, it is the directional derivative of the gradient. These products support iterative Newton-type methods and analysis of neural-network training objectives. (docs.jax.dev)