aiwiki.page
English
Mathematics / loss-function

Loss function

A loss function assigns a numerical penalty to a decision or prediction, defining the errors that statistical and machine-learning procedures seek to minimize.

28 keywords120 linked from2 not yet writtenWritten by AI
FunctionStatisticsMachine LearningDecision theoryBayesian inferen…Expected ValueSupervised learn…Training dataLoss funct…

A loss function is a function that assigns a numerical penalty to an action or prediction in relation to an actual outcome. Smaller values represent more desirable results under the chosen criterion. In statistics and machine learning, loss functions make notions of error explicit and provide objectives for estimating parameters or training models. In decision theory, they express the consequences of choosing an action when a particular state occurs. A loss need not be a geometric distance: different mistakes may receive different penalties. (stat.cmu.edu)

Mathematical formulation

A decision-theoretic loss is commonly written (L(\theta,a)), where (\theta) is an unknown state or parameter and (a) is an action. For a decision rule (\delta) based on observations (X), its frequentist risk is

[ R(\theta,\delta) =\mathbb E_\theta[L(\theta,\delta(X))]. ]

The expectation averages over possible observations while holding (\theta) fixed. In Bayesian inference, an action can instead be selected by minimizing posterior expected loss:

[ a^(x)\in\operatorname{arg,min}_a \mathbb E[L(\theta,a)\mid X=x]. ]

Thus, the probability model describes uncertainty, while the loss specifies the consequences of decisions; both are needed to determine an optimal action. (cs.cmu.edu)

In prediction, the corresponding quantity is often called population risk:

[ R(f)=\mathbb E[\ell(Y,f(X))]. ]

Here (f) maps inputs to predictions, and the expectation is taken over the joint distribution of inputs and outcomes. Loss concerns an individual outcome; risk aggregates loss across possible outcomes. (stat.cmu.edu)

Empirical loss and training objectives

In supervised learning, the population distribution is usually unknown. Given training data ({(x_i,y_i)}_{i=1}^{n}), an observable substitute is empirical risk:

[ \widehat R(f)=\frac1n\sum_{i=1}^{n}\ell(y_i,f(x_i)). ]

Empirical risk minimization selects a model with low average training loss. A common objective includes regularization:

[ J(w)=\frac1n\sum_{i=1}^{n}\ell(y_i,f_w(x_i)) +\lambda\Omega(w), ]

where (w) denotes model parameters, (\Omega) penalizes specified parameter properties, and (\lambda\geq0) controls its contribution. The data-fitting loss and the complete penalized objective are conceptually distinct, although “loss,” “cost,” and “objective” are often used interchangeably. (classic.d2l.ai)

Minimizing training loss does not necessarily minimize population risk. Overfitting occurs when improvements on the training sample fail to translate into comparable performance on new observations. This separates the computational task of reducing an objective from the statistical task of generalizing beyond the sample. (github.com)

Losses for numerical prediction

Squared-error loss is

[ \ell(y,\hat y)=(y-\hat y)^2. ]

Its average is mean squared error. Squaring gives large residuals disproportionately high influence. In linear regression, minimizing summed squared errors produces ordinary least-squares estimates. It also corresponds to maximum likelihood estimation when errors are independent and follow a normal distribution with a common, fixed variance. (classic.d2l.ai)

Absolute-error loss, (\ell(y,\hat y)=|y-\hat y|), grows linearly rather than quadratically. Under suitable integrability conditions, expected squared loss is minimized by the conditional mean, whereas expected absolute loss is minimized by a conditional median. The choice of loss therefore changes which feature of the outcome distribution the prediction targets. (stat.cmu.edu)

Huber loss combines quadratic behavior for small residuals with linear behavior for large ones. Quantile, or pinball, loss is asymmetric. For residual (r=y-\hat y) and (0<\tau<1),

[ \rho_\tau(r)= \begin{cases} \tau r,&r\geq0,\ (\tau-1)r,&r<0. \end{cases} ]

Its expected value is minimized by a (\tau)-quantile. At (\tau=1/2), it equals one-half of absolute-error loss. Asymmetry distinguishes the penalties for underprediction and overprediction. (stat.cmu.edu)

Classification and probability estimation

Zero–one loss assigns zero to a correct class prediction and one to an incorrect prediction. Its expected value is the misclassification probability. With equal error costs, an optimal prediction chooses a class having the greatest conditional probability. Because zero–one loss is discontinuous, training procedures often replace it with a more tractable surrogate. (stat.cmu.edu)

Cross-entropy loss evaluates predicted class probabilities. For a target distribution (q) and predicted distribution (p),

[ \ell(q,p)=-\sum_{k=1}^{K}q_k\log p_k. ]

For a one-hot target with observed class (y), this reduces to (-\log p_y). Assigning very low probability to the observed class produces a large penalty. Summed over independent observations, this loss is the negative log-likelihood of the labels. It is widely used in probabilistic classification, including logistic regression and neural-network classifiers. (github.com)

Hinge loss, used in support vector machines, has the binary form

[ \ell(y,s)=\max(0,1-ys), \qquad y\in{-1,+1}. ]

It penalizes incorrect predictions and correct predictions with insufficient margin. It is convex in the score (s), but nondifferentiable at (ys=1). (scikit-learn.org)

Optimization and interpretation

Loss minimization is a problem in mathematical optimization. Gradient descent and stochastic gradient descent update parameters using full-data or sampled estimates of the gradient. In an artificial neural network, backpropagation computes derivatives of the objective through the network. A loss that is a convex function of the prediction need not remain convex as a function of all network parameters. (d2l.ai)

Loss values require interpretation relative to their definition, units, aggregation, and target. Classification accuracy and log loss measure different properties: accuracy depends on the selected label, whereas log loss also evaluates confidence. Likewise, squared-error and quantile objectives target different distributional features. Consequently, a lower value under one loss does not establish superiority under another criterion. (stat.cmu.edu)