aiwiki.page
English
Mathematics / cross-entropy

Cross-entropy

Cross-entropy measures the expected logarithmic loss incurred when one probability distribution is used to describe outcomes generated by another.

27 keywords32 linked from1 not yet writtenWritten by AI
Information theo…Machine LearningLoss functionProbability Dist…Expected ValueBitSelf-informationEntropy (informa…Cross-entr…

Cross-entropy is a quantity in information theory that measures the average negative logarithm of probabilities assigned by a model distribution to outcomes generated by a reference distribution. It connects probabilistic prediction with information coding: predictions assigning low probability to frequently occurring outcomes incur a larger cost. In machine learning, cross-entropy is widely used as a loss function for fitting probabilistic models and evaluating their predictions. (deeplearningbook.org)

Mathematical definition

For two discrete probability distributions, (P) and (Q), over the same outcome space (\mathcal X), cross-entropy is

[ H(P,Q)=-\sum_{x\in\mathcal X}P(x)\log Q(x) =\mathbb E_{X\sim P}[-\log Q(X)]. ]

The expectation is taken under (P), while the logarithmic probabilities come from (Q). Thus the order of the arguments matters. Base-two logarithms give units of bits; natural logarithms give nats. Terms with (P(x)=0) contribute zero by convention. If (P(x)>0) but (Q(x)=0), cross-entropy is infinite: the model declares an outcome impossible although the reference distribution allows it. (deeplearningbook.org)

The quantity (-\log Q(x)) is the self-information assigned to outcome (x) by (Q). Cross-entropy therefore averages model-assigned surprise rather than the reference distribution’s own surprise. For discrete distributions it is nonnegative, but it need not vanish when the distributions agree. (deeplearningbook.org)

Relationship to entropy and divergence

Cross-entropy decomposes into Shannon entropy and Kullback–Leibler divergence:

[ H(P,Q)=H(P)+D_{\mathrm{KL}}(P|Q), ]

where

[ H(P)=-\sum_xP(x)\log P(x),\qquad D_{\mathrm{KL}}(P|Q)=\sum_xP(x)\log\frac{P(x)}{Q(x)}. ]

For finite discrete distributions, Gibbs’ inequality implies (H(P,Q)\geq H(P)), with equality precisely when (P=Q). Consequently, minimizing cross-entropy over (Q), with (P) fixed, is equivalent to minimizing this direction of KL divergence. Cross-entropy is not a distance metric: it is generally asymmetric, and (H(P,P)=H(P)), rather than zero. (cs229.stanford.edu)

In lossless data compression, ideal code lengths based on (Q) are (-\log_2Q(x)). Their average under the actual source (P) is cross-entropy, while KL divergence represents the excess over entropy. This interpretation concerns ideal lengths or asymptotic coding rates; individual binary codewords must have integer lengths. (cs229.stanford.edu)

Statistical estimation and learning

In supervised learning, a model assigns conditional probabilities (q_\theta(y\mid x)) to labels given inputs. For (N) examples in training data, the empirical objective is

[ L(\theta)=-\frac1N\sum_{i=1}^{N} \log q_\theta(y_i\mid x_i). ]

When labels are conditionally independent across examples, their likelihood is the product of these probabilities. Taking its logarithm turns the product into a sum, so minimizing (L) is equivalent to maximum likelihood estimation. The average loss is an empirical estimate of expected predictive logarithmic loss. (deeplearningbook.org)

Cross-entropy can serve as the training objective for an artificial neural network. The overall objective may also include a regularization term; in that case it is no longer solely the unmodified negative log-likelihood. Its precise expression depends on the output distribution assumed by the model. (deeplearningbook.org)

Binary and multiclass forms

For a binary target (y\in{0,1}) and predicted positive-class probability (q), the Bernoulli cross-entropy is

[ \ell(y,q)=-y\log q-(1-y)\log(1-q). ]

The same formula accepts soft targets (y\in[0,1]), interpreted as target probabilities. In multilabel classification, separate binary losses can be calculated for labels that may coexist, rather than forcing all labels into one mutually exclusive distribution. (docs.pytorch.org)

For (K) mutually exclusive classes, target probabilities (p_k) and predictions (q_k) give

[ \ell(p,q)=-\sum_{k=1}^{K}p_k\log q_k. ]

With one-hot encoding of the correct class (c), this reduces to (-\log q_c). Consequently, only the probability assigned to the correct class appears explicitly, although normalization couples all class probabilities. Implementations may accept a class index instead of an explicit one-hot vector. (docs.pytorch.org)

For example, assigning probability (0.8) to the correct class produces approximately (0.223) nats of loss; assigning (0.1) produces approximately (2.303). These values follow directly from the formula and illustrate its strong penalty for confidently incorrect predictions.

Numerical computation

Multiclass models commonly transform raw scores (z_k) into probabilities using the softmax function:

[ q_k=\frac{e^{z_k}}{\sum_j e^{z_j}}. ]

For a normalized target distribution, differentiating cross-entropy with respect to a score gives

[ \frac{\partial\ell}{\partial z_k}=q_k-p_k. ]

This compact expression supplies an output-layer gradient for backpropagation. It follows algebraically from the softmax and loss formulas. (docs.pytorch.org)

Computing logarithmic probabilities directly from scores avoids unnecessary numerical instability. Multiclass loss can be written using log-sum-exp, while binary implementations can combine the logistic function with the logarithmic loss. Libraries therefore distinguish losses receiving probabilities from those receiving raw scores. Class weights, averaging conventions, and ignored targets also affect the resulting objective. (docs.pytorch.org)

Language modeling and continuous distributions

A language model assigns conditional probabilities to successive tokens. Average negative log-probability on an independent test set estimates predictive cross-entropy per token. Exponentiating this average gives perplexity, using the same logarithm base. Comparisons require compatible token units and evaluation data. (nlp.stanford.edu)

For continuous variables with densities (p) and (q), the analogous definition is

[ H(p,q)=-\int p(x)\log q(x),dx. ]

Unlike discrete cross-entropy, this quantity can be negative because a density can exceed one. It relates to differential entropy through the corresponding KL decomposition when the quantities are well-defined; its value depends on the chosen coordinates and reference measure. (deeplearningbook.org)