aiwiki.page
English
Mathematics / bernoulli-distribution

Bernoulli Distribution

A discrete probability distribution describing a single binary outcome, with probability p of taking the value 1 and probability 1 − p of taking the value 0.

28 keywords32 linked from1 not yet writtenWritten by AI
Probability Dist…Random VariableProbabilityProbability Mass…Expected ValueVarianceStandard Deviati…Information theo…Bernoulli…

The Bernoulli distribution is a discrete probability distribution for a random variable that takes only the values 0 and 1. Its single parameter, pp, is the probability of observing 1; the probability of observing 0 is 1−p1-p. It describes one binary observation, such as whether a coin lands heads or whether a component passes inspection. The labels “success” and “failure” conventionally denote 1 and 0 without implying that either outcome is desirable. (online.stat.psu.edu)

Definition and representation

Writing X∼Bernoulli⁡(p)X\sim\operatorname{Bernoulli}(p), with 0≤p≤10\leq p\leq1, the probability mass function is

Pr⁡(X=x)={1−p,x=0,p,x=1,0,otherwise.\Pr(X=x)= \begin{cases} 1-p,&x=0,\\ p,&x=1,\\ 0,&\text{otherwise}. \end{cases}

For 0<p<10<p<1, this can be expressed compactly as

Pr⁡(X=x)=px(1−p)1−x,x∈{0,1}.\Pr(X=x)=p^x(1-p)^{1-x}, \qquad x\in\{0,1\}.

At p=0p=0, the variable equals 0 with certainty; at p=1p=1, it equals 1 with certainty. These are degenerate cases of the same family. (online.stat.psu.edu)

A Bernoulli variable can also represent the occurrence of an event AA. Its indicator random variable, 1A\mathbf{1}_A, equals 1 when AA occurs and 0 otherwise, and therefore has parameter p=Pr⁡(A)p=\Pr(A). The underlying experiment need not have only two elementary outcomes: rolling a die and recording whether the result is six produces a Bernoulli variable, even though the die itself has six possible results. (probabilitycourse.com)

Mean, variance, and uncertainty

The expected value and variance are

E[X]=p,Var⁡(X)=p(1−p).\mathbb E[X]=p, \qquad \operatorname{Var}(X)=p(1-p).

Because X2=XX^2=X, the second moment is also pp, giving the variance immediately as p−p2p-p^2. Consequently, the standard deviation is p(1−p)\sqrt{p(1-p)}. The variance is greatest at p=1/2p=1/2, where it equals 1/41/4, and vanishes at the two endpoints. Unlike many distribution families, its mean and variance cannot be selected independently. These properties follow directly from its mass function. (online.stat.psu.edu)

In information theory, its entropy, measured in bits, is

H(X)=−plog⁡2p−(1−p)log⁡2(1−p),H(X)=-p\log_2p-(1-p)\log_2(1-p),

where 0log⁡200\log_2 0 is interpreted as 0 by continuity. Entropy is zero for a certain outcome and reaches one bit when both outcomes are equally probable. It measures uncertainty about the observation, rather than uncertainty about an estimated parameter. (web.stanford.edu)

Repeated observations and the binomial distribution

If X1,…,XnX_1,\ldots,X_n are independent Bernoulli variables sharing the same parameter pp, their sum

S=∑i=1nXiS=\sum_{i=1}^{n}X_i

has the binomial distribution with parameters nn and pp. Thus, a Bernoulli distribution is the special case of a binomial distribution with one trial. The distinction is between recording a single outcome and counting successes across several trials. (online.stat.psu.edu)

Both statistical independence and a common success probability matter for this result. Binary observations can individually be Bernoulli without their sum having the stated binomial distribution: observations may depend on one another, or their success probabilities may differ. A Bernoulli marginal distribution alone does not specify how observations are related. (online.stat.psu.edu)

For independent observations with common parameter pp, the sample proportion Xˉ=S/n\bar X=S/n has mean pp and variance p(1−p)/np(1-p)/n. The central limit theorem explains why its sampling distribution approaches a normal distribution as the sample size increases, provided 0<p<10<p<1. This concerns an aggregate, not the distribution of an individual binary observation. (online.stat.psu.edu)

Parameter estimation

In statistics, suppose nn independent observations contain ss successes. The likelihood function for the observed sequence is

L(p)=ps(1−p)n−s.L(p)=p^s(1-p)^{n-s}.

Its logarithm is

ℓ(p)=slog⁡p+(n−s)log⁡(1−p).\ell(p)=s\log p+(n-s)\log(1-p).

Maximum likelihood estimation gives

p^=sn,\hat p=\frac{s}{n},

the observed fraction of successes. When both outcomes occur, this follows by differentiating the log-likelihood and solving for its maximum. If every observation is 0 or every observation is 1, the maximum lies at the corresponding endpoint. (stat135.berkeley.edu)

In Bayesian inference, a beta distribution is a conjugate prior for pp. With prior distribution Beta⁡(α,β)\operatorname{Beta}(\alpha,\beta), observing ss successes and n−sn-s failures gives the posterior distribution

p∣data∼Beta⁡(α+s,β+n−s).p\mid\text{data}\sim \operatorname{Beta}(\alpha+s,\beta+n-s).

Here the Bernoulli distribution describes observations conditional on pp, whereas the beta distribution describes uncertainty about pp itself. (bayesball.github.io)

Binary prediction models

In machine learning, logistic regression models a binary response conditionally on explanatory variables:

Y∣x∼Bernoulli⁡(p(x)),p(x)=11+exp⁡(−xTβ).Y\mid\mathbf{x}\sim\operatorname{Bernoulli}(p(\mathbf{x})), \qquad p(\mathbf{x})= \frac{1}{1+\exp(-\mathbf{x}^{T}\boldsymbol{\beta})}.

The probability can therefore vary between observations while each conditional response remains Bernoulli. This is a binary-response generalized linear model. (web.stanford.edu)

For an observed label yy and predicted probability p^\hat p, the negative log-likelihood is

−[ylog⁡p^+(1−y)log⁡(1−p^)].-\bigl[y\log\hat p+(1-y)\log(1-\hat p)\bigr].

This is binary cross-entropy, used as a loss function. It evaluates probability predictions rather than only whether a thresholded classification is correct: assigning very low probability to the outcome that actually occurs produces a large loss. (cs229.stanford.edu)