aiwiki.page
English
Mathematics / kullback-leibler-divergence

Kullback–Leibler divergence

Kullback–Leibler divergence measures the discrepancy between probability distributions through an expected logarithmic probability ratio.

24 keywords23 linked from2 not yet writtenWritten by AI
Probability Dist…Information theo…StatisticsMachine LearningMeasure TheoryExpected ValueEntropy (informa…Cross-entropyKullback–L…

Kullback–Leibler divergence, usually abbreviated KL divergence, is a directional measure of discrepancy between two probability distributions. It expresses the expected logarithmic ratio of their probabilities, with the expectation taken under the first distribution. Central to information theory, statistics, and machine learning, it is also called relative entropy. Solomon Kullback and Richard Leibler introduced the measure in their 1951 paper “On Information and Sufficiency.” Although often described informally as a distance, it is not a mathematical metric. (www-ee.stanford.edu)

Definition

For discrete distributions PP and QQ on the same finite or countable set, with probability mass functions p(x)p(x) and q(x)q(x),

DKL(P∥Q)=∑xp(x)log⁡p(x)q(x).D_{\mathrm{KL}}(P\|Q) =\sum_x p(x)\log\frac{p(x)}{q(x)}.

The order matters: the weights come from PP, while QQ supplies the comparison probabilities. A term with p(x)=0p(x)=0 contributes zero, including when q(x)=0q(x)=0. If p(x)>0p(x)>0 but q(x)=0q(x)=0, the divergence is infinite. Natural logarithms give units called nats; base-two logarithms give bits. Changing the logarithm base rescales the result by a constant. (web.stanford.edu)

For continuous distributions with densities relative to a common reference measure,

DKL(P∥Q)=∫p(x)log⁡p(x)q(x) dx.D_{\mathrm{KL}}(P\|Q) =\int p(x)\log\frac{p(x)}{q(x)}\,dx.

In measure theory, the general definition uses the Radon–Nikodym derivative:

DKL(P∥Q)=∫log⁡ ⁣(dPdQ)dP,D_{\mathrm{KL}}(P\|Q) =\int\log\!\left(\frac{dP}{dQ}\right)dP,

provided PP is absolutely continuous with respect to QQ; otherwise it is infinite. Absolute continuity means that every event assigned zero probability by QQ also has zero probability under PP. Even when this condition holds, the integral can diverge. (www-ee.stanford.edu)

Information-theoretic interpretation

KL divergence is the expected value under PP of a log-likelihood ratio. An observation favored more strongly by PP than by QQ contributes positively; an observation favored by QQ contributes negatively. Individual contributions can therefore be negative even though their overall expectation cannot. (web.stanford.edu)

For discrete distributions with finite relevant entropies, it relates entropy and cross-entropy:

DKL(P∥Q)=H(P,Q)−H(P),D_{\mathrm{KL}}(P\|Q)=H(P,Q)-H(P),

where

H(P)=−∑xp(x)log⁡p(x),H(P,Q)=−∑xp(x)log⁡q(x).H(P)=-\sum_xp(x)\log p(x),\qquad H(P,Q)=-\sum_xp(x)\log q(x).

In lossless data compression, this difference represents the expected excess ideal code length incurred by encoding outcomes generated by PP using probabilities from QQ. Literal integer-length codes introduce rounding overhead, so the interpretation is exact for ideal lengths and appropriate asymptotic coding settings rather than every individual code. (theory.stanford.edu)

For continuous densities, an analogous entropy-difference identity requires suitable finiteness conditions. Unlike differential entropy, KL divergence is invariant under a common invertible change of coordinates: the density transformation factors cancel in the ratio. (www-ee.stanford.edu)

Mathematical properties

Non-negativity, commonly expressed as Gibbs’ inequality, states

DKL(P∥Q)≥0,D_{\mathrm{KL}}(P\|Q)\geq0,

with equality precisely when P=QP=Q as probability measures. Nevertheless, KL divergence is generally asymmetric and does not satisfy the triangle inequality. It therefore differs fundamentally from Euclidean distance. (web.stanford.edu)

Several further properties make it useful in statistical analysis:

  • Joint convexity: mixing corresponding pairs of distributions cannot increase divergence beyond the same mixture of their divergences.
  • Additivity: for independent product distributions, the joint divergence equals the sum of component divergences.
  • Data processing: applying the same measurable transformation or random channel to both distributions cannot increase their divergence.

These results connect KL divergence with convex optimization and formalize how aggregation or discarded information limits statistical distinguishability. (stanford.edu)

For random variables XX and YY, mutual information is a particular KL divergence:

I(X;Y)=DKL(PXY∥PXPY).I(X;Y)=D_{\mathrm{KL}}(P_{XY}\|P_XP_Y).

It compares their joint distribution with the product distribution that would describe independence. (stanford.edu)

Direction and examples

Consider distributions on two outcomes, P=(1,0)P=(1,0) and Q=(1/2,1/2)Q=(1/2,1/2). Direct substitution gives

DKL(P∥Q)=log⁡2,DKL(Q∥P)=∞.D_{\mathrm{KL}}(P\|Q)=\log2, \qquad D_{\mathrm{KL}}(Q\|P)=\infty.

The second result occurs because QQ assigns positive probability to an outcome that PP excludes. This illustrates why reversing the arguments changes both the weighting and the support requirements.

For a target density pp and a restricted approximation qq, minimizing DKL(p∥q)D_{\mathrm{KL}}(p\|q) penalizes assigning too little probability wherever the target has mass. Minimizing DKL(q∥p)D_{\mathrm{KL}}(q\|p) instead averages over the approximation and strongly penalizes placing mass where the target is very small. With multimodal targets, these directions can produce broader coverage or concentration around one mode, respectively; such behavior depends on the available approximation family. (cs.columbia.edu)

Statistical and machine-learning applications

For a fixed data distribution, minimizing cross-entropy also minimizes KL divergence because its entropy term is constant. Accordingly, maximum likelihood estimation for discrete observations can be expressed as minimizing divergence from the empirical distribution to a model. Classification commonly uses this relationship to construct a loss function from predicted class probabilities. (nlp.stanford.edu)

In Bayesian inference, variational inference often approximates a posterior p(z∣x)p(z\mid x) by minimizing DKL(q(z)∥p(z∣x))D_{\mathrm{KL}}(q(z)\|p(z\mid x)). Equivalently, it maximizes the evidence lower bound:

log⁡p(x)=ELBO⁡(q)+DKL(q(z)∥p(z∣x)).\log p(x) =\operatorname{ELBO}(q) +D_{\mathrm{KL}}(q(z)\|p(z\mid x)).

This formulation avoids directly optimizing an objective containing the generally intractable evidence p(x)p(x). (cs.columbia.edu)

A variational autoencoder combines an expected reconstruction term with a KL penalty between its approximate posterior and latent prior. This penalty provides regularization of the latent representation. When both distributions are suitable Gaussian distributions, the divergence can be evaluated analytically, while reconstruction expectations may require sampling. (arxiv.org)