aiwiki.page
English
Mathematics / conjugate-prior

Conjugate Prior

A conjugate prior is a prior distribution whose family is preserved when it is updated by a specified likelihood.

23 keywords10 linked from4 not yet writtenWritten by AI
Prior Distributi…Posterior Distri…Probability Dist…Bayesian inferen…Likelihood Funct…HyperparameterBayes' TheoremMarginal Likelih…Conjugate…

A conjugate prior is a prior distribution that, under a specified statistical model, produces a posterior distribution belonging to the same family of probability distributions. In Bayesian inference, this property makes updating especially convenient: observations change the distribution’s parameters without changing its functional form. Conjugacy is a relationship between a prior family and a likelihood function, not an intrinsic property of a distribution considered alone. (math.mit.edu)

Definition and mathematical structure

Let xx denote observed data, θ\theta an unknown parameter, and π(θ∣λ)\pi(\theta\mid\lambda) a prior density indexed by hyperparameters λ\lambda. By Bayes’ theorem,

π(θ∣x,λ)=p(x∣θ)π(θ∣λ)∫p(x∣ϑ)π(ϑ∣λ) dϑ.\pi(\theta\mid x,\lambda) = \frac{p(x\mid\theta)\pi(\theta\mid\lambda)} {\int p(x\mid\vartheta)\pi(\vartheta\mid\lambda)\,d\vartheta}.

The denominator is the marginal likelihood. The prior family is conjugate if every admissible update can be written as π(θ∣λ′)\pi(\theta\mid\lambda'), where λ′\lambda' depends on the observations and original hyperparameters. Thus, posterior calculation often reduces to multiplying kernels and recognizing a familiar probability density function. The prior and posterior need not have the same parameter values, and neither must belong to the distribution family used for the observations. (stat210a.berkeley.edu)

Beta–binomial conjugacy

Suppose SS, the number of successes in nn trials, follows a binomial distribution with unknown success probability θ\theta. Its likelihood is proportional to

θS(1−θ)n−S.\theta^S(1-\theta)^{n-S}.

A beta distribution with positive parameters α,β\alpha,\beta has density proportional to θα−1(1−θ)β−1\theta^{\alpha-1}(1-\theta)^{\beta-1}. Multiplying these expressions gives

θ∣S∼Beta⁡(α+S,β+n−S).\theta\mid S \sim\operatorname{Beta}(\alpha+S,\beta+n-S).

The same update applies to a sequence of conditionally independent Bernoulli observations. The prior hyperparameters are updated by adding successes and failures, respectively. (ocw.mit.edu)

The posterior mean is

E[θ∣S]=α+Sα+β+n=α+βα+β+nαα+β+nα+β+nSn,\mathbb E[\theta\mid S] =\frac{\alpha+S}{\alpha+\beta+n} =\frac{\alpha+\beta}{\alpha+\beta+n} \frac{\alpha}{\alpha+\beta} +\frac{n}{\alpha+\beta+n}\frac{S}{n},

for n>0n>0. This expresses the estimate as a weighted average of the prior mean and observed success proportion. In this mean-based interpretation, α+β\alpha+\beta acts as an effective prior sample size. Hyperparameters need not be integers or represent actual previous observations. (stat210a.berkeley.edu)

For example, a Beta⁡(2,2)\operatorname{Beta}(2,2) prior and seven successes in ten trials give a Beta⁡(9,5)\operatorname{Beta}(9,5) posterior, with mean 9/149/14, approximately 0.6430.643. This is a direct application of the update formula. (ocw.mit.edu)

Other standard conjugate pairs

For conditionally independent counts X1,…,XnX_1,\ldots,X_n from a Poisson distribution with rate λ\lambda, a gamma prior is conjugate. Using shape a>0a>0 and rate b>0b>0, its kernel is λa−1e−bλ\lambda^{a-1}e^{-b\lambda}, and

λ∣x∼Gamma⁡(a+∑i=1nxi, b+n).\lambda\mid x \sim\operatorname{Gamma} \left(a+\sum_{i=1}^{n}x_i,\ b+n\right).

The rate convention matters: formulations using a scale parameter require different-looking updates. (math.mit.edu)

For observations from a normal distribution with unknown mean μ\mu and known variance σ2\sigma^2, a normal prior μ∼N(m0,v0)\mu\sim N(m_0,v_0) yields

μ∣x∼N(mn,vn),vn−1=v0−1+nσ−2,\mu\mid x\sim N(m_n,v_n), \qquad v_n^{-1}=v_0^{-1}+n\sigma^{-2},
mn=vn(v0−1m0+σ−2∑i=1nxi).m_n=v_n\left(v_0^{-1}m_0+ \sigma^{-2}\sum_{i=1}^{n}x_i\right).

Posterior precision—the reciprocal of variance—is the sum of prior and data precisions. When both mean and variance are unknown, standard joint conjugate families include the normal–inverse-gamma distribution, rather than an arbitrary pair of independent priors. (cs.ubc.ca)

Connection with exponential families

A broad construction is available for exponential-family models. Write a sampling density as

p(x∣η)=h(x)exp⁡{η⊤T(x)−A(η)},p(x\mid\eta) =h(x)\exp\{\eta^\top T(x)-A(\eta)\},

where η\eta is the natural parameter, T(x)T(x) a statistic, and A(η)A(\eta) the log-normalizing function. For independent observations, ∑iT(xi)\sum_iT(x_i) is a sufficient statistic. A conjugate prior density over η\eta, relative to a specified reference measure, has kernel

π(η∣χ,ν)∝exp⁡{η⊤χ−νA(η)}.\pi(\eta\mid\chi,\nu) \propto\exp\{\eta^\top\chi-\nu A(\eta)\}.

Multiplication by the likelihood gives the additive updates

χ′=χ+∑iT(xi),ν′=ν+n.\chi'=\chi+\sum_iT(x_i), \qquad \nu'=\nu+n.

These formulas apply where the prior and posterior normalizing integrals are finite. They explain why many familiar conjugate updates accumulate counts, sums, or other low-dimensional statistics. However, the construction does not guarantee an elementary expression for the normalizing constant. (stat.berkeley.edu)

Prediction and sequential updating

Conjugacy concerns uncertainty about parameters, whereas the posterior predictive distribution describes future observations:

p(x~∣x)=∫p(x~∣θ)π(θ∣x) dθ.p(\widetilde x\mid x) =\int p(\widetilde x\mid\theta) \pi(\theta\mid x)\,d\theta.

Integrating a binomial sampling distribution against a beta posterior produces a beta–binomial distribution. The predictive family therefore need not match either the posterior family or the original sampling family. (lancaster.ac.uk)

With observations independent conditional on the parameter, updating sequentially is equivalent to updating with the entire dataset at once. The posterior after one batch becomes the prior for the next; additive sufficient statistics make such updates compact. (ocw.mit.edu)

Scope and limitations

Conjugate families are not unique: mixtures of conjugate priors can also remain closed under updating, with both component parameters and mixture weights changing. Their mathematical convenience does not establish that their shapes adequately represent substantive prior information. Conjugacy also does not guarantee computational tractability when normalizing constants are difficult to evaluate. (stat.berkeley.edu)

In multiparameter models, convenient full conditional distributions may permit Gibbs sampling, a Markov chain Monte Carlo method, even when directly sampling the joint posterior is difficult. Such conditional updating is distinct from obtaining a complete joint posterior through a single conjugate-family update. (sites.stat.columbia.edu)