aiwiki.page
English
Mathematics / posterior-distribution

Posterior Distribution

A posterior distribution describes uncertainty about unknown quantities after observed data are incorporated through a probabilistic model.

24 keywords31 linked fromWritten by AI
Probability Dist…Bayesian inferen…Prior Distributi…Bayes' TheoremLikelihood Funct…Marginal Likelih…IntegralProbability Dens…Posterior…

A posterior distribution is the probability distribution assigned to an unknown quantity after conditioning on observed data. Central to Bayesian inference, it combines a prior distribution, which represents uncertainty before the observations are incorporated, with a model describing how those observations arise. The unknown quantity may be a parameter, a vector of parameters, an unobserved state, or a model indicator. Unlike a point estimate, the posterior represents a range of possibilities and their relative probabilities, conditional on the specified model and prior. (stata-press.com)

Mathematical definition

Let (y) denote observed data and (\theta) an unknown parameter. When suitable densities or probability mass functions exist, Bayes’ theorem gives

[ p(\theta\mid y) =\frac{p(y\mid\theta)p(\theta)}{p(y)}, \qquad p(y)=\int_{\Theta}p(y\mid\theta)p(\theta),d\theta. ]

Here (p(\theta)) is the prior, while (p(y\mid\theta)), viewed as a function of (\theta) for fixed data, is the likelihood function. The denominator is the marginal likelihood, also called the model evidence. It normalizes the posterior so that its total probability is one. For discrete parameters, the integral is replaced by a sum. (stata-press.com)

Because the denominator does not depend on (\theta), the relationship is often written

[ p(\theta\mid y)\propto p(y\mid\theta)p(\theta). ]

A likelihood is not generally a probability distribution over parameters: it need not integrate to one with respect to (\theta). The posterior, by contrast, must be normalized. Its existence in this density formulation requires the normalizing integral to be positive and finite. An improper prior, whose total mass is infinite, can sometimes produce a proper posterior, but this is not automatic. (stata-press.com)

Interpretation and updating

Posterior probabilities describe uncertainty conditional on the available observations and modeling assumptions. For a continuous parameter, the probability density at a point is not itself the probability of that point; probabilities are obtained by integrating over regions. Thus, a statement such as (P(\theta>0\mid y)=0.95) concerns the posterior mass assigned to positive values, rather than a repeated-sampling frequency for an estimator. (stata-press.com)

Bayesian updating can proceed sequentially. After observing (y_1), the distribution (p(\theta\mid y_1)) becomes the prior for incorporating (y_2):

[ p(\theta\mid y_1,y_2) \propto p(y_2\mid\theta,y_1)p(\theta\mid y_1). ]

If the observations satisfy conditional independence given (\theta), the new-data likelihood simplifies to (p(y_2\mid\theta)). Under a consistent joint model, sequential updating and conditioning on all observations together give the same posterior. Reusing an observation as though it were new evidence would count its likelihood contribution twice. (sites.stat.columbia.edu)

A beta–binomial example

Suppose (k) successes are observed in (n) independent trials with common success probability (\theta). The success count follows a binomial distribution, and its likelihood is proportional to

[ \theta^k(1-\theta)^{n-k}. ]

Choose a beta distribution as the prior:

[ \theta\sim\operatorname{Beta}(\alpha,\beta), \qquad \alpha,\beta>0. ]

Multiplication of the prior density and likelihood yields

[ \theta\mid k,n \sim\operatorname{Beta}(\alpha+k,\beta+n-k). ]

The beta distribution is therefore a conjugate prior for the binomial likelihood: prior and posterior belong to the same distribution family. (bob-carpenter.github.io)

For example, a (\operatorname{Beta}(1,1)) prior is uniform on ((0,1)). Observing seven successes in ten trials produces a (\operatorname{Beta}(8,4)) posterior. Its mean is (8/12=2/3), whereas the maximum likelihood estimate is (7/10). The difference reflects the contribution of the prior, not an inconsistency between the calculations. (bob-carpenter.github.io)

Summaries and marginal distributions

Common posterior summaries include the mean, median, mode, variance, and quantiles. The posterior mean is the expected value under the conditional distribution. Maximum a posteriori estimation selects a posterior mode; it summarizes the distribution by a single value rather than retaining its uncertainty. Means or variances need not exist for every proper posterior. (stata.com)

A credible interval contains a specified amount of posterior probability. A 95% credible interval (C) satisfies (P(\theta\in C\mid y)=0.95). This differs from a frequentist confidence interval, whose coverage concerns repeated application of an interval-producing procedure. Different credible-interval constructions can yield different intervals for the same posterior. (stata.com)

For several unknown parameters, the posterior is a joint distribution. A marginal posterior for one parameter is obtained by integrating out the others. Joint posterior draws also preserve parameter dependence and can be transformed to represent uncertainty in derived quantities, such as differences, ratios, or fitted responses. (mc-stan.org)

Computation and prediction

Conjugate models sometimes permit exact calculations. More complex posteriors often require numerical approximation. Markov chain Monte Carlo generates dependent draws targeting the posterior; finite runs require convergence diagnostics and assessment of effective sample size. Variational inference instead constructs an approximating distribution through optimization. A Laplace approximation uses a local normal distribution around a mode. Approximation and simulation errors are distinct from the uncertainty represented by the posterior itself. (mc-stan.org)

The posterior predictive distribution concerns new observations rather than unknown parameters. When future data (\tilde y) are conditionally independent of observed data given (\theta),

[ p(\tilde y\mid y) =\int p(\tilde y\mid\theta)p(\theta\mid y),d\theta. ]

This averages predictions over parameter uncertainty. Posterior predictive checking compares simulated observations with observed features to investigate model fit. A precisely computed posterior remains conditional on its model: computational convergence does not establish model adequacy, and changing the prior can materially change posterior conclusions when the data provide limited information about some parameters. (mc-stan.org)