The Jacobian matrix is a matrix containing the first-order partial derivatives of a vector-valued function. For a differentiable function, it represents the derivative as a linear map: multiplying a small input displacement by the Jacobian gives the first-order change in the output. It extends differentiation from single-variable functions to mappings between finite-dimensional spaces and connects multivariable calculus with linear algebra. (math.jhu.edu)
Definition and notation
Let , where is open, and write
If the relevant partial derivatives exist at , the Jacobian is
Thus has rows and columns: rows correspond to output components, and columns to input coordinates. Common notations include and , particularly when differentiability is established. (math.jhu.edu)
Using column vectors, a scalar-valued function has a Jacobian equal to the transpose of its gradient:
For , the Jacobian reduces to the ordinary derivative. For an affine mapping , it is the constant matrix . These conventions make derivative composition ordinary matrix multiplication. (ocw.mit.edu)
Differentiability and local approximation
The central interpretation is
The remainder divided by tends to zero. This requirement describes simultaneous changes in all input coordinates, not merely changes along coordinate axes. The Jacobian therefore represents the total derivative when this approximation holds. (math.cmu.edu)
Existence of all partial derivatives at a point does not by itself imply differentiability there. A standard sufficient condition is that the partial derivatives exist in a neighborhood and are continuous at the point. This distinction matters because a matrix of partial derivatives can exist even when it fails to provide a valid first-order approximation. (math.cmu.edu)
For example, consider the polynomial mapping
Direct differentiation gives
At , an input displacement produces
The linear terms are exactly ; the remaining terms are higher-order.
Composition and inverse mappings
For differentiable mappings and , the multivariable chain rule states
The order matters: the right-hand matrix first converts an input displacement into an intermediate displacement, and the left-hand matrix converts that into an output displacement. Their dimensions are respectively and . (live.ocw.mit.edu)
The inverse function theorem links a square Jacobian to local invertibility. If is continuously differentiable near and is invertible, then has a continuously differentiable inverse on suitable neighborhoods, with
The right-hand side is an inverse matrix. The conclusion is local, not a guarantee of global invertibility. A zero derivative does not automatically exclude an inverse: is invertible, although its inverse is not differentiable at zero. (math.colostate.edu)
Determinants and changes of variables
When , the determinant is called the Jacobian determinant. It must be distinguished from the matrix itself; rectangular Jacobians have no ordinary determinant. Its absolute value measures the local volume-scaling factor of a differentiable coordinate transformation, while its sign distinguishes preservation from reversal of orientation. (live.ocw.mit.edu)
For a continuously differentiable bijection between open subsets of , with continuously differentiable inverse, the change-of-variables formula for a suitable integrand is
The absolute value is essential because ordinary volume is unsigned. This generalizes substitution in a one-dimensional integral. (live.ocw.mit.edu)
For polar coordinates,
Its determinant is , yielding for . The formula is applied on regions where the angular coordinate makes the transformation one-to-one; the origin is a singular point of this coordinate system. (live.ocw.mit.edu)
Numerical computation and machine learning
For a nonlinear system , Newton’s method determines a correction from the linear system
This solves the local linear approximation rather than the original nonlinear equations directly. Implementations can solve for the correction without explicitly forming the inverse Jacobian. (web.mit.edu)
In automatic differentiation, the full matrix need not be stored. Forward-mode differentiation computes Jacobian–vector products ; reverse-mode computes transpose-Jacobian products . These propagate input perturbations and output sensitivities, respectively. Full Jacobians can be assembled column by column or row by row. (docs.jax.dev)
This distinction is important in machine learning. Backpropagation repeatedly applies reverse-mode products to obtain parameter gradients of a scalar loss function, avoiding explicit construction of every intermediate Jacobian. The Hessian matrix, by contrast, contains second derivatives of a scalar function; it can be viewed as the Jacobian of that function’s gradient. (live.ocw.mit.edu)