The law of large numbers is a family of theorems in probability describing the long-run behavior of averages of random variables. In its standard form, it states that the average of independent observations drawn from the same distribution converges to their expected value, provided that the expectation is finite. It supplies a mathematical connection between theoretical expectations and observed frequencies or averages. The weak and strong laws differ in the kind of convergence they establish, rather than in the numerical value of the limit. (ocw.mit.edu)
Mathematical formulation
Let (X_1,X_2,\ldots) be real-valued random variables defined on a common probability space. Assume statistical independence and that all observations have the same probability distribution; these assumptions are abbreviated i.i.d., meaning “independent and identically distributed.” Suppose [ \mathbb E[|X_1|]<\infty,\qquad \mu=\mathbb E[X_1]. ] The sample mean of the first (n) observations is [ \overline X_n=\frac{1}{n}\sum_{i=1}^{n}X_i. ] Both standard laws identify (\mu) as its limit. Finite variance is sufficient but is not required: finite absolute expectation already guarantees these i.i.d. results. (ocw.mit.edu)
The weak law states that, for every (\varepsilon>0), [ \lim_{n\to\infty} \Pr!\left(|\overline X_n-\mu|>\varepsilon\right)=0. ] This is convergence in probability. For any fixed error tolerance, the probability that an average lies outside that tolerance approaches zero. The assertion concerns probabilities at progressively larger sample sizes; it does not directly describe the entire trajectory of one infinite sequence of observations. (ocw.mit.edu)
The strong law states [ \Pr!\left(\lim_{n\to\infty}\overline X_n=\mu\right)=1. ] This is almost sure convergence. Except on an event of probability zero, each infinite observation sequence has averages converging to (\mu). Along such a sequence, every positive tolerance is eventually satisfied permanently, although the required sample size depends on the sequence. Almost sure convergence implies convergence in probability, so the strong law implies the weak law. Probability one nevertheless does not mean that every conceivable sequence converges. (ocw.mit.edu)
Relative frequencies and an elementary proof
For independent repetitions of an experiment with success probability (p), let (X_i=1) for success and (X_i=0) otherwise. Each observation has a Bernoulli distribution, with mean (p) and variance (p(1-p)). The average is the proportion of successes, while their total has a binomial distribution. The law therefore states that relative frequency converges to (p). For a fair coin, the proportion of heads converges to (1/2), not necessarily monotonically. (ocw.mit.edu)
When the common variance is finite, say (\sigma^2), independence gives [ \operatorname{Var}(\overline X_n)=\frac{\sigma^2}{n}. ] Chebyshev’s inequality then yields [ \Pr(|\overline X_n-\mu|\geq\varepsilon) \leq\frac{\sigma^2}{n\varepsilon^2}, ] which approaches zero. This provides a short proof of the weak law under the additional variance assumption. It also gives a finite-sample bound, although the bound need not be sharp. (ocw.mit.edu)
Strong-law proofs require additional control over deviations across an infinite sequence of sample sizes. One illustrative approach assumes a finite fourth moment, bounds the fourth moment of the average, and applies the Borel–Cantelli lemmas. That extra assumption simplifies the proof but is not part of the general i.i.d. theorem. (ocw.mit.edu)
Relationship to the central limit theorem
The law of large numbers identifies where averages converge. The central limit theorem instead describes the distribution of appropriately rescaled fluctuations. For i.i.d. observations with finite, positive variance, [ \frac{\sqrt n(\overline X_n-\mu)}{\sigma} \xrightarrow{d}N(0,1), ] where the arrow denotes convergence in distribution and (N(0,1)) is the standard normal distribution. (ocw.mit.edu)
Thus, in this finite-variance setting, the standard error of the sample mean is (\sigma/\sqrt n). Quadrupling the sample size halves this measure of variability. Normal approximations can support confidence intervals, whereas the law of large numbers alone supplies neither an approximate fluctuation distribution nor a universal sample size sufficient for a specified accuracy. (stat.berkeley.edu)
Statistical applications and limitations
In statistics, the theorem establishes that the sample mean is a consistent estimator of a finite population mean under i.i.d. sampling: its estimation error converges to zero in probability. Applying the same theorem to an integrable function (g(X_i)) also justifies estimating (\mathbb E[g(X)]) by an average of transformed observations. These are direct consequences of the theorem, with the assumptions applied to the quantities being averaged. (ocw.mit.edu)
In machine learning, an average loss for a fixed predictor can converge to its expected loss. A predictor selected from the same data presents a different issue: pointwise convergence for every fixed predictor does not automatically control the data-dependent selection. Empirical risk minimization consequently motivates uniform laws of large numbers, which control discrepancies simultaneously over a class of predictors. (stat.berkeley.edu)
The assumptions cannot simply be discarded. As an illustrative deduction, if every (X_i) equals the same nonconstant random variable (Y), then (\overline X_n=Y) for every (n); collecting more perfectly dependent observations does not force convergence to (\mathbb E[Y]). Nor does convergence require deviations to decrease at every step: the laws concern limiting behavior, not monotonic improvement or exact equality after finitely many observations. (ocw.mit.edu)
Historical development
Jacob Bernoulli proved an early weak law for repeated success-or-failure trials in Ars Conjectandi, published posthumously in 1713. His theorem quantified how sufficiently many repetitions make observed proportions likely to approximate their underlying probability. Subsequent work broadened the theorem beyond Bernoulli trials; twentieth-century results included weak laws requiring only a finite first absolute moment in the i.i.d. setting. (mathshistory.st-andrews.ac.uk)