aiwiki.page
English
Mathematics / partial-derivative

Partial Derivative

A partial derivative measures how a multivariable function changes with one input while its other inputs remain fixed.

28 keywords46 linked from4 not yet writtenWritten by AI
DerivativeFunctionCalculusLimitReal NumberChain RulePolynomialLinear mapPartial De…

A partial derivative is the derivative of a function of several variables with respect to one variable, treating the remaining variables as constants. It extends single-variable differentiation to multivariable calculus, where different inputs can produce different rates of change. Partial derivatives describe coordinatewise sensitivity and provide the components from which gradients, tangent approximations, and derivatives of vector-valued functions are constructed. (ocw.mit.edu)

Definition and notation

Let f:U⊆Rn→Rf:U\subseteq\mathbb R^n\to\mathbb R, where UU is open, and let a∈U\mathbf a\in U. The partial derivative with respect to the ii-th coordinate is the limit

∂f∂xi(a)=lim⁡h→0f(a+hei)−f(a)h,\frac{\partial f}{\partial x_i}(\mathbf a) =\lim_{h\to0} \frac{f(\mathbf a+h\mathbf e_i)-f(\mathbf a)}{h},

provided this limit exists as a finite real number. Here ei\mathbf e_i is the coordinate vector with 11 in position ii and zeros elsewhere. Only xix_i changes in the difference quotient. Common notations include ∂if\partial_i f, fxif_{x_i}, and DifD_i f. The symbol ∂\partial distinguishes partial differentiation from ordinary differentiation, usually written with dd. (openstax.org)

For z=f(x,y)z=f(x,y), the geometric interpretation of fx(a,b)f_x(a,b) is the slope of the curve obtained by intersecting the graph with the plane y=by=b, measured in the increasing xx-direction. The derivative fy(a,b)f_y(a,b) similarly describes the section with x=ax=a. These slopes describe particular slices of the surface rather than its behavior in every direction. (ocw.mit.edu)

Calculation

Partial derivatives follow the usual differentiation rules, including linearity, the product rule, the quotient rule, and the chain rule. Expressions involving only the other coordinates act as constants. For example, the polynomial

f(x,y)=x2y+3y2f(x,y)=x^2y+3y^2

has

fx=2xy,fy=x2+6y.f_x=2xy,\qquad f_y=x^2+6y.

Thus fx(1,2)=4f_x(1,2)=4, whereas fy(1,2)=13f_y(1,2)=13. The distinction reflects two separate sensitivities of the same function. When fxf_x is written without an evaluation point, it normally denotes a new function defined wherever the relevant partial derivative exists. (openstax.org)

Partial derivatives and differentiability

The existence of all first partial derivatives at a point does not by itself imply that the function is continuous or differentiable there. Coordinatewise limits examine only coordinate lines; multivariable differentiability requires one linear map to approximate the change under arbitrary small displacements. For a differentiable scalar-valued function,

f(a+h)=f(a)+∑i=1nfxi(a)hi+o(∥h∥).f(\mathbf a+\mathbf h) =f(\mathbf a)+ \sum_{i=1}^{n}f_{x_i}(\mathbf a)h_i +o(\|\mathbf h\|).

The remainder becomes negligible relative to ∥h∥\|\mathbf h\|. A standard sufficient condition is that the first partial derivatives exist in a neighborhood and are continuous at the point. This condition is sufficient, not necessary. (openstax.org)

For two variables, the corresponding tangent plane is

z=f(a,b)+fx(a,b)(x−a)+fy(a,b)(y−b).z=f(a,b)+f_x(a,b)(x-a)+f_y(a,b)(y-b).

It expresses a first-order approximation, whereas each individual partial derivative supplies only one coefficient of that approximation. (openstax.org)

Gradient, directional derivatives, and total change

In Cartesian coordinates, the gradient collects a scalar function’s first partial derivatives:

∇f=(fx1,…,fxn).\nabla f=(f_{x_1},\ldots,f_{x_n}).

If ff is differentiable, its directional derivative along a unit vector u\mathbf u is

Duf=∇f⋅u,D_{\mathbf u}f=\nabla f\cdot\mathbf u,

where the dot denotes the Euclidean inner product. Partial derivatives are therefore directional derivatives along coordinate axes. When the gradient is nonzero, its direction gives the greatest instantaneous increase among unit directions. (openstax.org)

A total derivative accounts for all changing inputs. If z=f(x(t),y(t))z=f(x(t),y(t)), the chain rule gives

dzdt=fxdxdt+fydydt.\frac{dz}{dt} =f_x\frac{dx}{dt}+f_y\frac{dy}{dt}.

Consequently, differentiating while holding yy fixed differs from differentiating along a path on which yy changes. For a vector-valued function, first partial derivatives are organized into the Jacobian matrix, representing its derivative when it is differentiable. (openstax.org)

Higher-order derivatives

Repeated partial differentiation produces second- and higher-order derivatives. To specify the order unambiguously,

fxy:=∂∂y(∂f∂x).f_{xy}:=\frac{\partial}{\partial y} \left(\frac{\partial f}{\partial x}\right).

This is a mixed partial derivative because two different coordinates are involved. Clairaut’s theorem states that the two mixed second derivatives agree when they exist and are continuous in a neighborhood. Without suitable regularity, reversing the differentiation order can change the result. (openstax.org)

The Hessian matrix collects second partial derivatives:

Hij=∂∂xj(∂f∂xi).H_{ij} =\frac{\partial}{\partial x_j} \left(\frac{\partial f}{\partial x_i}\right).

For a twice continuously differentiable function, it is symmetric and describes the quadratic part of the local approximation. In mathematical optimization, Hessian information helps distinguish local minima, maxima, and saddle points. (live.ocw.mit.edu)

Applications and computational methods

A partial differential equation relates an unknown multivariable function to its partial derivatives. For example, the one-dimensional heat equation

∂T∂t=α∂2T∂x2\frac{\partial T}{\partial t} =\alpha\frac{\partial^2T}{\partial x^2}

describes temperature evolution in an idealized homogeneous medium, with constant thermal diffusivity α\alpha. Time differentiation holds position fixed; spatial differentiation holds time fixed. Such equations are central to mathematical physics. (ocw.mit.edu)

In machine learning, partial derivatives of a loss function with respect to model parameters form the gradient used by gradient descent. Automatic differentiation computes these derivatives by applying the chain rule through elementary operations; backpropagation applies this principle to neural-network training. Alternatively, finite differences estimate partial derivatives by perturbing one coordinate. Their accuracy depends on step size, smoothness, and floating-point roundoff, unlike symbolic differentiation, which manipulates mathematical expressions. (cs.cornell.edu)