An estimator is a rule in statistics for using sample data to infer an unknown quantity, such as a population mean, variance, or model parameter. It is usually expressed as a function of the observations. An estimator must be distinguished from an estimate: the estimator is the rule, whereas an estimate is the particular value obtained by applying that rule to observed data. Because different samples generally produce different values, an estimator has statistical properties that can be studied under repeated sampling. (ocw.mit.edu)
Mathematical definition
Let denote data drawn from a statistical model whose probability distribution depends on an unknown parameter . An estimator of is written
The rule depends on the data and known model information, not on the unknown parameter itself. Before observation, is a random variable; after observing , its realized value is the estimate. The distribution of under repeated sampling is its sampling distribution. (ocw.mit.edu)
A point estimator returns a single value, or a vector when several parameters are estimated jointly. An interval-estimation procedure instead returns a range. A confidence interval is evaluated through its coverage probability over repeated samples, rather than solely through the distance between a point estimate and the parameter. (ocw.mit.edu)
Basic examples
For independent, identically distributed observations with finite mean and variance , the sample mean
estimates . Its expected value is , and its variance is . Thus, under these assumptions, increasing the sample size reduces its sampling variability. (ocw.mit.edu)
For , the usual sample-variance estimator is
It is unbiased for . The denominator , rather than , implements Bessel’s correction. For samples from a normal distribution with unknown mean, maximum likelihood instead produces the variance estimator with denominator , illustrating that different estimation criteria can yield different rules for the same quantity. (ocw.mit.edu)
Bias, error, and consistency
The bias of an estimator is
An estimator is unbiased if this difference is zero for every parameter value in the model. Unbiasedness describes an average over possible samples; it does not guarantee that any individual estimate is close to the truth. Moreover, an unbiased estimator of a parameter need not remain unbiased after a nonlinear transformation. (ocw.mit.edu)
For a scalar parameter, mean squared error measures expected squared estimation error:
Consequently, a biased estimator can outperform an unbiased estimator under squared-error loss if its variance is sufficiently smaller. This distinguishes minimizing total error from merely eliminating bias. (ocw.mit.edu)
A sequence is a consistent estimator if it approaches in probability:
Consistency concerns increasing sample size, whereas unbiasedness is a finite-sample expectation property. Neither generally implies the other. The law of large numbers establishes consistency of the sample mean under suitable assumptions. (hsong1.github.io)
Methods of construction
The method of moments constructs estimators by equating empirical moments with corresponding model moments and solving for the unknown parameters. For example, because a uniform distribution on has mean , moment matching gives . (its.caltech.edu)
Maximum likelihood estimation chooses a parameter value maximizing the likelihood function for the observed sample. For independent observations from a Bernoulli distribution, the maximum likelihood estimator of the success probability is the sample proportion. Likelihood maximizers need not always exist or be unique. (its.caltech.edu)
In Bayesian inference, a prior distribution and the likelihood determine a posterior distribution. A Bayesian point estimator minimizes posterior expected loss. Under squared-error loss, it is the posterior mean; under absolute-error loss, a posterior median is optimal. The choice of loss therefore helps determine what constitutes an appropriate estimate. (ocw.mit.edu)
Efficiency and information
Among unbiased estimators of the same parameter, efficiency is often assessed by variance. Under suitable regularity conditions, the Cramér–Rao bound gives
where is the sample’s Fisher information. An unbiased estimator attaining this bound is efficient in this sense; the bound need not be attainable in every model. (ocw.mit.edu)
A sufficient statistic retains all information in the sample about the parameter within the specified model. The Rao–Blackwell theorem provides an improvement procedure: replacing a finite-variance estimator by its conditional expectation given a sufficient statistic preserves its expectation and cannot increase its mean squared error. (ocw.mit.edu)
Sampling uncertainty
An estimator’s standard error is the standard deviation of its sampling distribution, often itself estimated from data. For the sample mean under independent sampling, it is . Standard error measures uncertainty in the estimated mean, not the spread of individual observations. (ocw.mit.edu)
The central limit theorem supports normal approximations for many estimators, including sample means with finite, nonzero variance. Such approximations can support confidence intervals, but their validity depends on the sampling assumptions and sample size. Estimation theory accordingly separates the rule producing an estimate from the assumptions used to characterize its uncertainty. (hsong1.github.io)