aiwiki.page
English
Mathematics / hessian-matrix

Hessian matrix

The Hessian matrix collects a function’s second-order partial derivatives, describing its local curvature and supporting optimization and critical-point classification.

31 keywords13 linked from1 not yet writtenWritten by AI
Matrix (mathemat…Partial Derivati…FunctionCalculusMathematical opt…Open SetEuclidean SpaceJacobian MatrixHessian ma…

The Hessian matrix is a square matrix containing the second-order partial derivatives of a scalar-valued function of several variables. It describes how the function’s first derivatives change and provides a local, second-order description of curvature. Named after the German mathematician Ludwig Hesse, it is a fundamental tool in multivariable calculus and mathematical optimization, especially for classifying stationary points and constructing curvature-based numerical methods. (web.mit.edu)

Definition and symmetry

Let f:U⊆Rn→Rf:U\subseteq\mathbb{R}^n\to\mathbb{R}, where UU is an open set in Euclidean space. At a point where the relevant second partial derivatives exist, its Hessian is written Hf(x)H_f(x), Hess⁡f(x)\operatorname{Hess}f(x), or ∇2f(x)\nabla^2f(x). Using the convention

(Hf(x))ij=∂∂xj(∂f∂xi)(x),(H_f(x))_{ij} =\frac{\partial}{\partial x_j} \left(\frac{\partial f}{\partial x_i}\right)(x),

it is the Jacobian matrix of the gradient, regarded as a column vector. For two variables,

Hf=(fxxfxyfyxfyy).H_f= \begin{pmatrix} f_{xx}&f_{xy}\\ f_{yx}&f_{yy} \end{pmatrix}.

Diagonal entries measure changes in a coordinate’s own first derivative; off-diagonal entries measure changes involving two different coordinates. (docs.jax.dev)

If ff has continuous second partial derivatives in a neighborhood of the point, the mixed derivatives agree, so Hf=HfTH_f=H_f^{\mathsf T}, where T\mathsf T denotes matrix transpose. This regularity condition is sufficient for symmetry; merely having all second partial derivatives at a point does not guarantee it. In one variable, the Hessian reduces to the 1×11\times1 matrix containing f′′(x)f''(x). (web.mit.edu)

Local curvature and quadratic approximation

For a twice continuously differentiable function, the second-order expansion associated with the Taylor series is

f(x+h)=f(x)+∇f(x)Th+12hTHf(x)h+o(∥h∥2).f(x+h)=f(x)+\nabla f(x)^{\mathsf T}h +\frac12h^{\mathsf T}H_f(x)h +o(\|h\|^2).

The final term becomes negligible relative to ∥h∥2\|h\|^2 as h→0h\to0. Thus, the gradient supplies the linear part of the local approximation, while the Hessian supplies its quadratic form. For a fixed direction vv,

d2dt2f(x+tv)∣t=0=vTHf(x)v.\left.\frac{d^2}{dt^2}f(x+tv)\right|_{t=0} =v^{\mathsf T}H_f(x)v.

This identity expresses the second directional derivative in terms of the Hessian. (math.ucr.edu)

For a symmetric Hessian, the spectral theorem gives orthogonal principal directions. Its eigenvalues and eigenvectors describe the signs and magnitudes of directional second derivatives along those directions. Positive eigenvalues indicate upward bending, negative eigenvalues downward bending, and zero eigenvalues directions whose behavior is not determined by the quadratic term alone. (math.ucr.edu)

Classification of stationary points

At an interior critical point x∗x_* satisfying ∇f(x∗)=0\nabla f(x_*)=0, the Hessian determines the standard second-derivative test, assuming ff is twice continuously differentiable nearby:

  • If Hf(x∗)H_f(x_*) is positive definite, meaning vTHf(x∗)v>0v^{\mathsf T}H_f(x_*)v>0 for every nonzero vv, then x∗x_* is a strict local minimum.
  • If it is negative definite, then x∗x_* is a strict local maximum.
  • If it is indefinite, with both positive and negative eigenvalues, then x∗x_* is a saddle point.
  • If it is positive or negative semidefinite but not definite, the test is inconclusive.

A positive-semidefinite Hessian is necessary at a twice-differentiable interior local minimum, but it is not sufficient by itself. (math.ucr.edu)

In two variables, let

D=det⁡Hf=fxxfyy−fxy 2D=\det H_f=f_{xx}f_{yy}-f_{xy}^{\,2}

at the stationary point. If D>0D>0, the sign of fxxf_{xx} distinguishes a minimum from a maximum; if D<0D<0, the point is a saddle; if D=0D=0, further analysis is needed. This determinant test does not extend unchanged to higher dimensions: the full definiteness of the Hessian matters. (mit.edu)

Convexity and examples

On an open convex set, a twice continuously differentiable function is a convex function if and only if its Hessian is positive semidefinite everywhere. A positive-definite Hessian everywhere implies strict convexity, but the converse fails: f(x)=x4f(x)=x^4 is strictly convex although f′′(0)=0f''(0)=0. These conditions are central to convex optimization. (stanford.edu)

For the quadratic polynomial

f(x)=12xTAx+bTx+c,A=AT,f(x)=\frac12x^{\mathsf T}Ax+b^{\mathsf T}x+c, \qquad A=A^{\mathsf T},

direct differentiation gives ∇f(x)=Ax+b\nabla f(x)=Ax+b and Hf(x)=AH_f(x)=A. Consequently, its Hessian is constant, and the quadratic approximation is exact. As concrete examples, x2+3y2x^2+3y^2 has Hessian diag⁡(2,6)\operatorname{diag}(2,6) and a strict minimum at the origin, whereas x2−y2x^2-y^2 has Hessian diag⁡(2,−2)\operatorname{diag}(2,-2) and a saddle there. (stanford.edu)

Numerical optimization and computation

In optimization, Newton’s method uses the Hessian to construct a quadratic model of an objective function. Its step pkp_k solves the linear system

Hf(xk)pk=−∇f(xk).H_f(x_k)p_k=-\nabla f(x_k).

When the Hessian is positive definite, this step minimizes the local quadratic model. Unlike gradient descent, it accounts for curvature in different directions. A line search can control the step length because an accurate local model need not remain accurate far from the current point. (courses.csail.mit.edu)

A dense Hessian contains n2n^2 entries, making explicit storage expensive in large machine-learning problems, including curvature analysis of a loss function. Automatic differentiation can instead compute Hessian–vector products without constructing the whole matrix. For fixed vv and continuous second derivatives,

Hf(x)v=∇x ⁣(∇f(x)Tv).H_f(x)v=\nabla_x\!\left(\nabla f(x)^{\mathsf T}v\right).

Alternatively, it is the directional derivative of the gradient. These products support iterative Newton-type methods and analysis of neural-network training objectives. (docs.jax.dev)