aiwiki.page
English
Mathematics / prior-distribution

Prior Distribution

A prior distribution represents uncertainty about an unknown quantity before the data used in a particular Bayesian analysis are incorporated.

27 keywords35 linked from4 not yet writtenWritten by AI
Probability Dist…Bayesian inferen…Posterior Distri…Likelihood Funct…Bayes' TheoremMarginal Likelih…IntegralProbability Dens…Prior Dist…

A prior distribution is a probability distribution assigned to an unknown quantity before incorporating the observations used in a particular analysis. In Bayesian inference, it combines with a model for the observations to produce a posterior distribution. Priors can express information from previous studies, scientific constraints, or assumptions about plausible parameter values. They can also stabilize estimation when observations provide limited information. “Prior” is relative to the data being analyzed: it does not necessarily mean that the distribution was specified before any relevant evidence existed. (stat.columbia.edu)

Mathematical role

Let θ\theta denote an unknown parameter or parameter vector, and let yy denote observed data. Write the prior density as p(θ)p(\theta) and the likelihood function as p(y∣θ)p(y\mid\theta). Bayes’ theorem gives

p(θ∣y)=p(y∣θ)p(θ)∫p(y∣ϑ)p(ϑ) dϑ∝p(y∣θ)p(θ).p(\theta\mid y) = \frac{p(y\mid\theta)p(\theta)} {\int p(y\mid\vartheta)p(\vartheta)\,d\vartheta} \propto p(y\mid\theta)p(\theta).

The denominator is the marginal likelihood, or evidence. It normalizes the posterior, provided it is finite and positive. For discrete parameters, a sum replaces the integral. The prior describes uncertainty over parameter values; the likelihood describes how those values account for the observed data. A likelihood is not generally a normalized distribution over parameters. (web.stanford.edu)

A prior may be specified through a probability density function, probabilities on discrete alternatives, or a joint distribution over several unknowns. Its influence depends on its interaction with the likelihood, rather than on its spread considered in isolation. (arxiv.org)

Information and parameterization

An informative prior incorporates substantive information about an unknown quantity. A weakly informative prior imposes broad restrictions, such as a plausible order of magnitude, without tightly concentrating uncertainty around a particular value. A prior described as noninformative or objective instead follows a formal rule intended to limit particular kinds of prior influence. These labels are contextual, not universal measures of information. (arxiv.org)

Uniformity does not establish complete ignorance: a density uniform in one parameterization is generally not uniform after a nonlinear transformation. Formal constructions such as Jeffreys’ prior address invariance under reparameterization. Even a very diffuse prior can substantially affect inference through the parameter scale, the likelihood, and the predictions implied by their combination. (arxiv.org)

Conjugate priors and an example

A conjugate prior belongs to a family that remains unchanged in form after updating with a specified likelihood. Conjugacy makes posterior calculations analytically tractable; it is a relationship between a prior family and a likelihood, not an intrinsic property of a distribution alone. (web.stanford.edu)

For example, suppose ss successes are observed in nn independent trials with success probability θ\theta. A binomial likelihood combined with a beta prior gives

θ∼Beta⁡(α,β),θ∣s,n∼Beta⁡(α+s,β+n−s),\theta\sim\operatorname{Beta}(\alpha,\beta), \qquad \theta\mid s,n\sim \operatorname{Beta}(\alpha+s,\beta+n-s),

where α,β>0\alpha,\beta>0. The posterior mean is

E[θ∣s,n]=α+sα+β+n.E[\theta\mid s,n]=\frac{\alpha+s}{\alpha+\beta+n}.

For n>0n>0, this is a weighted average of the prior mean and the observed success proportion; α+β\alpha+\beta controls their relative weighting. (web.stanford.edu)

As an illustrative calculation, a Beta⁡(2,2)\operatorname{Beta}(2,2) prior and eight successes in ten trials yield a Beta⁡(10,4)\operatorname{Beta}(10,4) posterior. Its mean is 10/1410/14, approximately 0.7140.714, compared with the maximum-likelihood estimate of 0.80.8. The posterior remains a distribution rather than merely an adjusted point estimate. (web.stanford.edu)

Hierarchical specification

In a hierarchical Bayesian model, related parameters share a distribution governed by hyperparameters. For example, group-specific effects might follow a normal distribution with an unknown population mean and variance. Assigning priors to these hyperparameters allows uncertainty about the population distribution to be represented alongside uncertainty about individual groups. This structure distinguishes variation between groups from uncertainty about their overall level. (stat.columbia.edu)

Proper and improper priors

A proper prior has total probability one. An improper prior, such as a constant density over the entire real line, cannot be normalized and is therefore not literally a probability distribution. Nevertheless, its product with a likelihood can sometimes produce a proper posterior. Posterior propriety depends on the model and observations; it is not automatic. (stat.columbia.edu)

Improper priors pose additional problems in model comparison because their arbitrary multiplicative constants leave ordinary marginal likelihoods undetermined. Bayes factors, which compare marginal likelihoods, therefore require special treatment when improper priors are involved. Even proper but highly diffuse priors can strongly influence model comparisons. (arxiv.org)

Regularization and model assessment

In machine learning and statistics, priors can implement regularization. Maximum a posteriori estimation minimizes a negative log-likelihood plus a negative log-prior. A Gaussian coefficient prior therefore produces a quadratic penalty; under Gaussian-error linear regression, this corresponds to ridge regression. Full Bayesian inference additionally propagates uncertainty rather than retaining only the posterior mode. (arxiv.org)

A prior predictive distribution describes data implied jointly by the prior and sampling model:

p(y)=∫p(y∣θ)p(θ) dθ.p(y)=\int p(y\mid\theta)p(\theta)\,d\theta.

Prior predictive checks simulate parameters from the prior, then observations conditional on those parameters. Implausible simulated outcomes can reveal unsuitable scales, locations, or distributional shapes. These checks differ from posterior predictive checking, which conditions on observed data. (mc-stan.org)

Sensitivity analysis compares posterior results under alternative plausible priors. Prior influence is often greater with sparse observations or weakly identified parameters, while well-identified parameters with sufficiently informative data may show little sensitivity. Its extent must be assessed within the particular model and inferential question. (stat.columbia.edu)