A directional derivative measures how a function changes at a particular point when its input moves along a specified direction. In multivariable calculus, it extends the ordinary derivative by restricting a function to a line through that point. For a real-valued function, the result is a scalar rate of change; when the direction vector has unit length, it measures change per unit distance. Coordinate-axis directions give the familiar partial derivatives. (ocw.mit.edu)
Definition and conventions
Let , where is an open subset of Euclidean space . At a point , the directional derivative along a vector is the limit
provided it exists as a finite number. Equivalently, defining reduces the question to whether the single-variable derivative exists. Only values on this particular line enter the definition. (math.ucdavis.edu)
Two conventions are common. Elementary geometric treatments usually require a unit vector , so the parameter represents signed distance. More general treatments allow arbitrary vectors, with their lengths specifying the speed of the parametrization. For nonzero , writing gives
whenever either derivative exists. Thus normalization changes the numerical rate, although not the direction of travel. Here denotes the Euclidean norm. (ocw.mit.edu)
Relation to differentiability and the gradient
If is differentiable at , it has a first-order approximation
where is a linear map. Substituting shows that . For scalar-valued functions, the gradient represents this map through the Euclidean inner product:
In particular, choosing a standard coordinate vector recovers the corresponding partial derivative. Differentiability is essential to this general formula: the mere existence of coordinate partial derivatives does not establish a linear approximation valid in all directions. (ocw.mit.edu)
The formula also follows from the chain rule. More generally, if a differentiable curve satisfies , then
Thus the instantaneous change along a curved path depends on its velocity at the point, rather than its subsequent course. (ocw.mit.edu)
Geometric interpretation and example
For a function of two variables, restricting its graph to a vertical plane in direction produces a cross-sectional curve. The directional derivative is that curve’s tangent slope. When , the largest directional derivative over unit vectors is , attained in the gradient’s direction; the smallest is its negative, attained in the opposite direction. Directions perpendicular to the gradient give zero first-order change. These directions are tangent to regular level curves or level surfaces. (ocw.mit.edu)
For example, consider the polynomial
At , its gradient is . Along the unit vector , direct substitution into the gradient formula gives
Along the unnormalized vector , the derivative is . The fivefold difference reflects the fivefold speed of compared with . This calculation illustrates the distinction between rate per unit distance and rate per parameter increment. (ocw.mit.edu)
Directional derivatives without differentiability
Existence of directional derivatives in every direction does not imply differentiability, or even continuity. A counterexample is
For a fixed direction with ,
If , the quotient is identically zero. Hence every directional derivative at the origin vanishes. Nevertheless, along the curved path , the function equals for , so it is not a continuous function at the origin. Direction-by-direction limits therefore need not describe behavior throughout a neighborhood. (math.ucdavis.edu)
Optimization and computation
In mathematical optimization, a negative directional derivative of a differentiable objective function guarantees decrease for sufficiently small positive steps along that vector. This underlies gradient descent: when the gradient is nonzero,
A line search chooses the step length separately; local derivative information alone does not determine how far a step should go. (cs.cornell.edu)
For a differentiable vector-valued function , the corresponding derivative is
where is its Jacobian matrix. Forward-mode automatic differentiation computes such Jacobian–vector products by propagating directional changes through successive operations. This permits evaluation of a specified directional response without explicitly constructing the entire Jacobian, including in calculations involving machine-learning models. (docs.jax.dev)