The chain rule is a theorem of calculus that determines the derivative of a composition of functions. It describes how changes propagate through successive dependencies: an input changes an intermediate quantity, which changes the final output. For scalar functions, the relevant derivatives multiply; for functions of several variables, contributions from different dependencies must also be added. The rule connects elementary differentiation with multivariable analysis and computational methods for evaluating derivatives. (openstax.org)
Single-variable formulation
Let be differentiable at , and let be differentiable at , with the composition defined near . Then
The outer derivative must be evaluated at the output of the inner function, not at the original input. These differentiability assumptions are sufficient; continuity of the derivative functions is not required. (openstax.org)
Using the notation associated with Gottfried Wilhelm Leibniz, if and , the same identity becomes
Although this resembles cancellation of fractions, its justification is the differentiation theorem. (live.ocw.mit.edu)
For example, taking a polynomial as the inner function,
For three nested functions, repeated application gives
Every layer contributes a derivative evaluated at its own input. Omitting an inner derivative is therefore an incomplete application of the rule. (openstax.org)
Mathematical basis
The rule follows from differentiability understood as local linear approximation. Put . For small ,
and, for small ,
Here denotes an error whose ratio to approaches zero in the relevant limit. Substituting the first increment into the second expansion produces
which establishes the derivative formula. This argument remains valid when , unlike a naive proof that divides by the change in , which may vanish. (openstax.org)
The assumptions are not necessary for every composite to be differentiable. For instance, is not differentiable at zero, but gives . Thus, a differentiable composite does not imply differentiable component functions. This example illustrates the distinction between sufficient hypotheses and their converse. (openstax.org)
Several variables
Suppose , with and , and assume the functions are differentiable. Then
Each partial derivative measures sensitivity to one intermediate variable, holding the other fixed. Both terms contribute because both intermediate variables can change with . The partial derivatives are evaluated at . (openstax.org)
For example, if , , and , then
consistent with . More generally, if and each depends on ,
Dependency diagrams organize this calculation by multiplying derivatives along each path and adding the resulting contributions. (openstax.org)
Jacobian and gradient forms
For differentiable maps and , the derivative is a linear map between vector spaces. The chain rule states
In coordinates, it becomes a product of Jacobian matrices:
The dimensions are respectively , , and . The order matters: matrix multiplication generally does not commute. This expresses the composition of local linear approximations in linear algebra. (live.ocw.mit.edu)
When is scalar-valued and gradients are column vectors,
The transpose converts the outer gradient into sensitivity with respect to the original input coordinates. (live.ocw.mit.edu)
Integration by substitution
The chain rule underlies integration by substitution. If , then
so
This is the differentiation identity used in reverse. Together with the fundamental theorem of calculus, it also gives the corresponding definite integral:
provided is continuously differentiable on and is continuous on its image. The integration bounds change with the substituted variable. (openstax.org)
Computational applications
Automatic differentiation evaluates derivatives of computer programs by decomposing calculations into elementary operations and repeatedly applying the chain rule. A computational graph records dependencies. Forward mode propagates input sensitivities toward outputs; reverse mode propagates output sensitivities backward. Unlike finite-difference approximations, these methods differentiate the operations themselves, although numerical evaluation remains subject to rounding. (jmlr.org)
In machine learning, backpropagation applies reverse-mode differentiation to obtain derivatives of a loss function with respect to parameters of an artificial neural network. Those derivatives support optimization methods such as gradient descent. When an intermediate value influences the output through several branches, its derivative contributions must be accumulated, reflecting the sum in the multivariable chain rule rather than a single product along one path. (jmlr.org)