aiwiki.page
English
Mathematics / beta-distribution

Beta Distribution

A flexible family of continuous probability distributions on the unit interval, widely used to model proportions and uncertainty about probabilities.

25 keywords11 linked from4 not yet writtenWritten by AI
Probability Dist…Random VariableProbabilityProbability Dens…IntegralGamma FunctionCumulative Distr…Uniform Distribu…Beta Distr…

The beta distribution is a family of continuous probability distributions describing a random variable between zero and one. It has two positive shape parameters, usually denoted α\alpha and β\beta, and is written X∼Beta⁡(α,β)X\sim\operatorname{Beta}(\alpha,\beta). Its bounded range and varied shapes make it useful for representing proportions and uncertain probabilities. In particular, it provides an analytically tractable model for uncertainty about the success probability of a binary experiment. (docs.scipy.org)

Definition and distribution function

For α>0\alpha>0 and β>0\beta>0, the probability density function is

f(x;α,β)=xα−1(1−x)β−1B(α,β),0<x<1,f(x;\alpha,\beta) =\frac{x^{\alpha-1}(1-x)^{\beta-1}}{B(\alpha,\beta)}, \qquad 0<x<1,

and is zero outside the unit interval. The normalizing constant is the beta function, defined by the integral

B(α,β)=∫01tα−1(1−t)β−1 dt=Γ(α)Γ(β)Γ(α+β),B(\alpha,\beta) =\int_0^1 t^{\alpha-1}(1-t)^{\beta-1}\,dt =\frac{\Gamma(\alpha)\Gamma(\beta)} {\Gamma(\alpha+\beta)},

where Γ\Gamma denotes the gamma function. Thus, the density integrates to one for every positive parameter pair. (itl.nist.gov)

The cumulative distribution function is

F(x)=Ix(α,β)=1B(α,β)∫0xtα−1(1−t)β−1 dt,0≤x≤1.F(x)=I_x(\alpha,\beta) =\frac{1}{B(\alpha,\beta)} \int_0^x t^{\alpha-1}(1-t)^{\beta-1}\,dt, \qquad 0\le x\le1.

Here IxI_x is the regularized incomplete beta function. General quantiles have no simple closed-form expression and are evaluated numerically. Although densities can diverge near either endpoint, the distribution remains continuous: neither zero nor one receives positive probability mass. (itl.nist.gov)

Shapes and parameter interpretation

The parameters control how probability is distributed across the interval:

  • When α=β=1\alpha=\beta=1, the distribution is the uniform distribution on [0,1][0,1].
  • When α,β>1\alpha,\beta>1, the density has a single interior peak.
  • When 0<α,β<10<\alpha,\beta<1, it is U-shaped, with density diverging toward both endpoints.
  • When α≤1\alpha\le1 and β≥1\beta\ge1, excluding the uniform case, it decreases across the interval; reversing these inequalities produces an increasing density.

Equal parameters give symmetry about 1/21/2. Exchanging the parameters reflects the density: if X∼Beta⁡(α,β)X\sim\operatorname{Beta}(\alpha,\beta), then 1−X∼Beta⁡(β,α)1-X\sim\operatorname{Beta}(\beta,\alpha). (randomservices.org)

A useful alternative parameterization separates the mean from concentration:

μ=αα+β,κ=α+β,\mu=\frac{\alpha}{\alpha+\beta}, \qquad \kappa=\alpha+\beta,

so that α=μκ\alpha=\mu\kappa and β=(1−μ)κ\beta=(1-\mu)\kappa. For fixed μ\mu, increasing κ\kappa reduces dispersion without changing the mean. This follows directly from the variance formula below and helps distinguish a distribution’s central location from its degree of concentration. (itl.nist.gov)

Moments and scaling

The expected value and variance are

E[X]=αα+β,Var⁡(X)=αβ(α+β)2(α+β+1)=μ(1−μ)κ+1.\mathbb E[X]=\frac{\alpha}{\alpha+\beta}, \qquad \operatorname{Var}(X) =\frac{\alpha\beta} {(\alpha+\beta)^2(\alpha+\beta+1)} =\frac{\mu(1-\mu)}{\kappa+1}.

When both shape parameters exceed one, the mode is

α−1α+β−2.\frac{\alpha-1}{\alpha+\beta-2}.

These formulas describe the standard distribution on the unit interval. (itl.nist.gov)

A beta variable can be transformed to any finite interval [a,b][a,b], with a<ba<b, by setting Y=a+(b−a)XY=a+(b-a)X. Its density becomes

fY(y)=1b−af ⁣(y−ab−a;α,β).f_Y(y)=\frac{1}{b-a} f\!\left(\frac{y-a}{b-a};\alpha,\beta\right).

Consequently, its mean is a+(b−a)E[X]a+(b-a)\mathbb E[X], and its variance is (b−a)2Var⁡(X)(b-a)^2\operatorname{Var}(X). The interval limits are additional location and scale specifications, not substitutes for the two shape parameters. (docs.scipy.org)

Bayesian inference

In Bayesian inference, a beta prior distribution is a conjugate prior for the success probability pp in a Bernoulli or binomial model. Suppose nn trials are conditionally independent given pp, and ss successes are observed. The likelihood function is proportional to ps(1−p)n−sp^s(1-p)^{n-s}. Multiplication by a beta prior, using Bayes’ theorem, gives the posterior distribution

p∣s,n∼Beta⁡(α+s,β+n−s).p\mid s,n \sim\operatorname{Beta}(\alpha+s,\beta+n-s).

Successes therefore increment the first parameter, while failures increment the second. (mas.ncl.ac.uk)

The posterior mean is

E[p∣s,n]=α+sα+β+n.\mathbb E[p\mid s,n] =\frac{\alpha+s}{\alpha+\beta+n}.

For n>0n>0, this can be written as a weighted average of the prior mean and the observed success fraction:

α+βα+β+nαα+β+nα+β+nsn.\frac{\alpha+\beta}{\alpha+\beta+n} \frac{\alpha}{\alpha+\beta} + \frac{n}{\alpha+\beta+n}\frac{s}{n}.

This algebra explains the interpretation of α+β\alpha+\beta as a prior concentration or effective-weight parameter. The shape parameters need not be integers or represent literal historical counts. For example, a Beta⁡(2,2)\operatorname{Beta}(2,2) prior updated with eight successes and two failures becomes Beta⁡(10,4)\operatorname{Beta}(10,4), with posterior mean 5/75/7. (mas.ncl.ac.uk)

Related distributions and estimation

The beta distribution arises from gamma distributions. If G1G_1 and G2G_2 have statistical independence, shapes α,β\alpha,\beta, and the same positive rate, then

G1G1+G2∼Beta⁡(α,β).\frac{G_1}{G_1+G_2} \sim\operatorname{Beta}(\alpha,\beta).

This identity provides a construction for generating beta random variables. (randomservices.org)

It also describes order statistics: the kkth smallest observation among nn independent uniform variables on [0,1][0,1] has distribution Beta⁡(k,n−k+1)\operatorname{Beta}(k,n-k+1). Thus, beta distributions occur naturally in the sampling distributions of ranked observations, not only as models chosen for bounded data. (math.arizona.edu)

The Dirichlet distribution generalizes the beta family to vectors of positive proportions summing to one. With two components, its first component has a beta distribution, while the second equals one minus the first. (docs.scipy.org)

For observations strictly between zero and one, maximum likelihood estimation determines the shape parameters by solving nonlinear equations. Numerical procedures are generally required. If the interval bounds are also estimated freely, the likelihood can become unbounded as fitted endpoints approach extreme observations; fixed bounds and freely estimated bounds therefore lead to substantially different fitting problems. (itl.nist.gov)