Bessel’s correction is an adjustment in statistics that replaces the denominator with when estimating a population’s variance from observations centered on their sample mean. For independent, identically distributed observations with finite variance, this adjustment produces an unbiased estimator of population variance. It compensates for the downward bias introduced by estimating the population mean from the same observations used to measure dispersion. The correction also appears in the conventional sample standard deviation, although it does not make that quantity unbiased. (numpy.org)
Definition and assumptions
Let , with , be random variables drawn from the same probability distribution, with statistical independence, common mean , and finite variance . Define the sample mean and squared-deviation sum by
The uncorrected variance estimator is ; the corrected estimator is
Its expected value satisfies
Unbiasedness concerns the average over repeated samples, not whether an individual estimate equals the population variance. These results do not require a normal distribution. (cs.cmu.edu)
The distinction between random variables and observed data is often reflected in notation: denotes the random estimator, while denotes its value after observation. Some texts call both and “sample variance,” so the denominator must be established from context. (statproofbook.github.io)
Why the correction works
A direct proof follows from the identity
The first term has expectation . Because the sample mean is unbiased and has variance , the second term has expectation . Consequently,
Dividing by therefore removes the bias exactly. The subtraction identifies the variation lost when the unknown population mean is replaced by a mean fitted to the observations. (cs.cmu.edu)
The same adjustment has a degrees-of-freedom interpretation. The deviations satisfy
so once deviations are specified, the final deviation is determined. In geometric terms, subtracting the mean is an orthogonal projection onto the linear subspace of vectors whose coordinates sum to zero. This subspace has dimension , rather than . The expectation calculation supplies the statistical justification; the constraint supplies its geometric interpretation. (web.stanford.edu)
Numerical example
Consider the illustrative observations . Their mean is , and their squared-deviation sum is
The uncorrected variance is , whereas the corrected sample variance is . The corresponding corrected standard deviation is . This example illustrates the formulas, not a guarantee that the corrected value is closer to an unknown population variance. (online.stat.psu.edu)
The correction multiplies the uncorrected variance by . It doubles it when , increases it by approximately 11.1% when , and becomes proportionally smaller as sample size increases. At , the corrected estimator is undefined: centering a single observation on itself leaves no residual degree of freedom. (numpy.org)
Scope and limitations
If the population mean is known, then
is already unbiased for . The adjustment is needed because the center is estimated from the same sample, not simply because the observations form a sample. Likewise, when describing an entire finite dataset rather than estimating a larger population’s variance, division by expresses its average squared deviation. (cs.cmu.edu)
For standard deviation, taking a square root introduces another bias. By Jensen’s inequality,
Thus, Bessel’s correction makes variance unbiased but generally leaves standard deviation biased downward. Unbiasedness also does not establish minimum mean squared error. Under normal sampling with unknown mean, maximum likelihood estimation yields , despite its downward bias. (numpy.org)
Related statistical applications
Under independent normal sampling, the sampling distribution obeys
This chi-squared distribution underlies variance tests and confidence intervals. Normality is needed for this exact distributional result, but not for the unbiasedness of . (itl.nist.gov)
In linear regression fitted by ordinary least squares, the analogous adjustment divides the residual sum of squares by , where counts all fitted coefficients, including an intercept. Under a correctly specified, full-rank model with zero-mean uncorrelated errors of common variance, this estimates error variance without bias. Estimating only a common mean is the special case . (stat.cmu.edu)
Software conventions can differ. NumPy’s var uses denominator by default; setting ddof=1 changes it to . Applying the same option to std returns the square root of the corrected variance, not an unbiased standard-deviation estimator. (numpy.org)