The evidence lower bound (ELBO) is an objective function used in variational inference to approximate otherwise intractable probability distributions. It is a lower bound on the logarithm of a model’s marginal likelihood, also called its evidence. Maximizing the ELBO makes an approximate distribution closer to the exact posterior in a specific divergence, while avoiding direct evaluation of the evidence. It is central to approximate Bayesian inference and to learning latent-variable models in machine learning. (cs.columbia.edu)
Mathematical definition
Let denote observed data, unobserved variables, and model parameters. A joint probability distribution is written
where is the prior distribution and supplies the likelihood. Computing the evidence requires marginalizing :
For discrete variables, the integral is replaced by a sum. The exact posterior distribution is . These quantities may be computationally difficult even when the joint density is evaluable. (cs.columbia.edu)
For an approximate distribution , define
Here denotes an expectation under . The defining relation is
Because Kullback–Leibler divergence is nonnegative, . Equality holds precisely when the approximate and exact posterior distributions agree almost everywhere. The identity assumes a positive, finite evidence and appropriately defined expectations; incompatible supports can instead yield an ELBO of negative infinity. (cs.columbia.edu)
Derivation and equivalent forms
The decomposition follows by substituting the posterior into the divergence:
Rearranging produces the ELBO identity. Alternatively, when covers the joint distribution’s support in , Jensen’s inequality applied to the concave logarithm gives
The left side is then . (cs.columbia.edu)
Two equivalent expressions highlight different interpretations:
and
The first combines expected log joint density with entropy; for continuous variables, is differential entropy. The second balances data fit against departure from the prior. Its negative is commonly called variational free energy, although sign conventions differ across fields. (arxiv.org)
Optimization and expectation–maximization
With fixed, maximizing the ELBO over a family is equivalent to minimizing . Thus inference becomes mathematical optimization. A restricted family can leave a nonzero gap even at its global optimum. (jmlr.org)
A common choice is a factorized, or mean-field, family . Coordinate-ascent updates optimize one factor while holding the others fixed:
The constant normalizes the factor. Such updates improve the objective when performed exactly, but generally need not find its global maximum. (jmlr.org)
The expectation–maximization algorithm admits the same variational interpretation. Its E-step sets to the current exact posterior, making the bound tight. Its M-step increases the bound with respect to , thereby increasing or preserving the observed-data log likelihood. Variational EM instead restricts , allowing approximate E-steps when exact inference is unavailable. Increasing a loose bound alone does not guarantee an increase in the true evidence. (cs.toronto.edu)
Neural models and stochastic computation
In a variational autoencoder, an inference network supplies , while a generative network defines . Both parameter sets are optimized through the ELBO. The expected log likelihood is often called the reconstruction term, while the prior divergence acts as regularization. Minimizing the negative ELBO expresses the same optimization as a loss function. (arxiv.org)
The reparameterization trick allows differentiable sampling constructions, such as
where denotes elementwise multiplication. This permits gradient estimates through sampled latent variables. Automatic differentiation and stochastic gradient methods support scalable optimization; stochastic variational inference also uses data subsampling to fit models with large datasets. (arxiv.org)
Interpretation and limitations
The exact ELBO is a bound, but a finite-sample stochastic estimate can exceed the log evidence because of sampling variability. Numerical optimization therefore distinguishes the mathematical objective from its noisy estimates. (arxiv.org)
The approximation also depends on the divergence’s direction. Minimizing can favor concentration in a posterior mode; factorization can suppress dependencies and underestimate uncertainty. A high ELBO is consequently not, by itself, proof of accurate posterior uncertainty. (cs.columbia.edu)
Importance-weighted objectives extend the construction by averaging several importance weights inside a logarithm. Their expected bounds are no looser than the single-sample ELBO and can better accommodate complex posteriors. They remain distinct objectives: tighter evidence bounds and improved optimization behavior are separate questions. (arxiv.org)