In statistics, the bias of an estimator is the difference between the estimator’s expected value and the true value of the quantity being estimated. It measures systematic overestimation or underestimation across repeated samples, rather than the error in one particular result. Bias is a property of an estimation procedure under a specified statistical model or sampling design; an unbiased procedure need not produce an accurate estimate in every sample. (stat.berkeley.edu)
Definition and interpretation
Let denote data drawn from a probability distribution indexed by an unknown parameter . An estimator is a function of the data used to estimate . Before the data are observed, it is a random variable with a sampling distribution. Its bias is
provided the expectation exists. The subscript indicates that the expectation is calculated under the distribution corresponding to . An estimator is unbiased if for every parameter value in the specified model. Being unbiased at one particular value does not establish unbiasedness throughout the model. (statlect.com)
Positive bias means that the estimator is too large on average; negative bias means that it is too small. The realized estimation error admits the decomposition
The second term has expectation zero and represents random sampling fluctuation. Thus, unbiasedness concerns the center of the sampling distribution, not its spread or the closeness of any individual estimate to the target. (stat.berkeley.edu)
Elementary examples
Suppose are identically distributed observations with finite mean . By linearity of expectation, the sample mean
satisfies , so it is unbiased for . If the observations also have statistical independence and common finite variance , then . This distinguishes unbiasedness from precision: the expectation identifies the target, whereas the variance describes sampling variability. (homepages.ecs.vuw.ac.nz)
A standard biased estimator is the sample variance calculated with denominator . For independent, identically distributed observations with finite variance and , define
Then
Replacing by , a modification called Bessel’s correction, produces an unbiased estimator of . Under a normal distribution with unknown mean, is nevertheless the maximum-likelihood estimator of the variance. Maximum likelihood and unbiasedness therefore need not select the same estimator. (stats-web.quarto.pub)
Bias, variance, and estimation loss
Bias alone does not determine an estimator’s quality. Under squared-error loss, the mean squared error is
This identity holds when the relevant second moments are finite. It follows by expanding the error around ; the cross term vanishes because the centered estimator has expectation zero. For an unbiased estimator, MSE equals variance. A biased estimator can have lower MSE if its reduction in variance outweighs its squared bias. (bookdown.org)
For example, let follow a binomial distribution with parameters and . The sample proportion is unbiased. The smoothed estimator
has, by direct calculation,
It pulls estimates toward and reduces variance, but its MSE advantage depends on . At , its bias vanishes and its variance is smaller than that of . (stat135.berkeley.edu)
This bias–variance tradeoff also appears in machine learning. In linear regression, ridge regression uses regularization to shrink coefficient estimates. Such shrinkage generally introduces bias but can reduce variability and improve prediction relative to an unpenalized fit. The objective is not necessarily exact unbiasedness, but lower expected prediction error under the chosen loss. (stat.cmu.edu)
Finite-sample bias and consistency
For a sequence , asymptotic unbiasedness means that approaches zero as sample size increases. The variance estimator above has this property despite being biased at every finite when . Consistency, by contrast, means convergence in probability to the target:
It describes concentration near the parameter, not merely convergence of the expectation. (stats-web.quarto.pub)
Unbiasedness does not imply consistency: an estimator that uses only the first observation can remain unbiased for the population mean while retaining the same sampling distribution as more observations become available. A useful sufficient condition for consistency is that both bias and variance approach zero. Then MSE approaches zero, and Chebyshev’s inequality applied to squared estimation error establishes convergence in probability. (mcrovella.github.io)
Estimating and correcting bias
When an expectation is known analytically, bias may be removable exactly. For example, if for a known nonzero constant , then is unbiased. Such corrections depend on the model and need not minimize MSE. (stat135.berkeley.edu)
Bootstrap sampling can approximate bias when an exact calculation is unavailable. If denotes the estimator computed from bootstrap data, the conditional difference serves as a bootstrap bias estimate. Subtracting it gives . Its reliability depends on how well the bootstrap distribution approximates the actual sampling distribution; estimating a correction does not guarantee exact unbiasedness. (stat.cmu.edu)