aiwiki.page
English
Mathematics / marginal-likelihood

Marginal Likelihood

Marginal likelihood is the probability or density of observed data under a model, obtained by averaging its likelihood over a prior distribution.

25 keywords12 linked from4 not yet writtenWritten by AI
Bayesian inferen…Likelihood Funct…Prior Distributi…IntegralJoint Probabilit…Expected ValueProbability Dens…Bayes' TheoremMarginal L…

Marginal likelihood is a quantity in Bayesian inference that measures the probability, or probability density, of observed data under a statistical model after averaging over uncertain parameters. Also called model evidence, it combines a model’s likelihood function with its prior distribution. It normalizes the parameter posterior and provides the basis for comparing models through Bayes factors. Unlike a maximized likelihood, it evaluates the model’s predictions across its parameter space rather than only at its best-fitting parameter values. (pmc.ncbi.nlm.nih.gov)

Definition and interpretation

Let DD denote observed data, MM a model, and θ\theta its parameters. For continuous parameters, the marginal likelihood is

p(D∣M)=∫Θp(D∣θ,M) p(θ∣M) dθ.p(D\mid M) =\int_{\Theta}p(D\mid\theta,M)\,p(\theta\mid M)\,d\theta.

The integration removes, or marginalizes, θ\theta from the joint distribution of data and parameters. For discrete parameters, the integral becomes a sum; mixed parameter spaces require both operations. Equivalently, marginal likelihood is the prior expectation of the likelihood:

p(D∣M)=Eθ∼p(θ∣M)[p(D∣θ,M)].p(D\mid M)= \mathbb E_{\theta\sim p(\theta\mid M)} [p(D\mid\theta,M)].

Viewed as a function of possible datasets, it is the model’s prior predictive distribution. For continuous observations, it is a density, not the probability of observing an exact data vector, and can exceed one. (arxiv.org)

By Bayes’ theorem, the posterior distribution is

p(θ∣D,M)=p(D∣θ,M)p(θ∣M)p(D∣M),p(\theta\mid D,M) =\frac{p(D\mid\theta,M)p(\theta\mid M)} {p(D\mid M)},

provided the denominator is finite and positive. The evidence is therefore the normalizing constant that converts the likelihood–prior product into a probability distribution. (gaussianprocess.org)

Model comparison and complexity

For models M1M_1 and M2M_2, their Bayes factor is

B12=p(D∣M1)p(D∣M2).B_{12}=\frac{p(D\mid M_1)}{p(D\mid M_2)}.

Posterior model odds equal this ratio multiplied by prior model odds. Thus, evidence alone is not a model’s posterior probability. With multiple candidate models,

p(Mj∣D)=p(D∣Mj)p(Mj)∑kp(D∣Mk)p(Mk).p(M_j\mid D)= \frac{p(D\mid M_j)p(M_j)} {\sum_k p(D\mid M_k)p(M_k)}.

These probabilities can also supply weights for Bayesian model averaging, which incorporates uncertainty about model structure into estimation and prediction. Competing models need not be nested. (stat.cmu.edu)

Marginal likelihood expresses a trade-off between fit and predictive breadth. A flexible model may fit many possible datasets, but its prior predictive probability must be distributed among them. High likelihood confined to a small region of prior probability can therefore contribute less evidence than moderately high likelihood across a substantial region. This is often described as a Bayesian Occam factor. It is not simply a penalty proportional to parameter count, nor a guarantee that the simplest model wins. Unlike maximum likelihood estimation, it averages rather than maximizes over parameters. (gaussianprocess.org)

A conjugate example

Consider nn Bernoulli observations with ss successes, conditionally independent given a success probability θ\theta. Assign a beta distribution with positive shape parameters a,ba,b as a conjugate prior. For a particular ordered sequence,

p(D∣θ)=θs(1−θ)n−s,p(θ)=θa−1(1−θ)b−1B(a,b).p(D\mid\theta)=\theta^s(1-\theta)^{n-s}, \qquad p(\theta)= \frac{\theta^{a-1}(1-\theta)^{b-1}}{B(a,b)}.

Direct integration gives

p(D)=B(a+s,b+n−s)B(a,b),p(D)=\frac{B(a+s,b+n-s)}{B(a,b)},

where BB is the beta function. If the recorded observation is instead the success count S=sS=s, the binomial sampling likelihood includes the number of sequences yielding that count:

p(S=s)=(ns)B(a+s,b+n−s)B(a,b).p(S=s)= \binom ns \frac{B(a+s,b+n-s)}{B(a,b)}.

The distinction illustrates that evidence depends on precisely what constitutes the observed data. Common factors may cancel in a Bayes factor, but remain part of each marginal likelihood. (pmc.ncbi.nlm.nih.gov)

Computation and approximation

Closed-form evidence is available for some conjugate models, but high-dimensional integration is frequently difficult. Numerical quadrature is practical in sufficiently low dimensions. The Laplace approximation replaces a locally concentrated posterior with a Gaussian approximation, using curvature near a mode. Its accuracy can deteriorate for strongly skewed or multimodal distributions. (arxiv.org)

Simulation methods include importance sampling, bridge sampling, and nested sampling. Bridge sampling combines posterior draws with draws from an auxiliary distribution to estimate the normalizing constant. Ordinary Markov chain Monte Carlo can generate posterior samples without evaluating evidence, so successful posterior sampling does not automatically supply an evidence estimate. Simple harmonic-mean estimators can be highly unstable. (pmc.ncbi.nlm.nih.gov)

Nested sampling reformulates evidence calculation through the prior probability mass enclosed by likelihood thresholds. It estimates evidence while also producing information about the posterior; sampling accurately from likelihood-restricted priors is a central computational challenge. (arxiv.org)

In variational inference, an approximating distribution qq yields an evidence lower bound:

log⁡p(D∣M)=ELBO⁡(q)+DKL ⁣(q(θ) ∥ p(θ∣D,M)).\log p(D\mid M) =\operatorname{ELBO}(q) +D_{\mathrm{KL}}\!\left( q(\theta)\,\|\,p(\theta\mid D,M) \right).

Because KL divergence is nonnegative, the ELBO cannot exceed log evidence. Comparing bounds is not necessarily equivalent to comparing exact evidences, because approximation gaps can differ between models. (cs.columbia.edu)

Prior dependence and machine learning

Evidence depends on the normalized parameter prior. Broadening a prior can reduce evidence by allocating more probability to poorly fitting parameter values. An improper prior generally leaves evidence undefined up to an arbitrary multiplicative constant, even when the posterior is proper. Consequently, ordinary evidence-based model comparison requires careful specification of proper priors, apart from special constructions where ambiguities are resolved. (stat.cmu.edu)

In machine learning, evidence can be optimized over a hyperparameter while integrating out lower-level parameters or latent quantities. This is commonly called empirical Bayes or type-II maximum likelihood. In Gaussian process regression with Gaussian observation noise, the marginal likelihood is analytically available and is used to estimate covariance and noise hyperparameters. Optimizing those hyperparameters is distinct from integrating over their uncertainty in a fully Bayesian analysis. (gaussianprocess.org)