In statistics, the standard error (SE) of a statistic is the standard deviation of its sampling distribution: the distribution of values that would arise from repeatedly drawing samples under the same sampling procedure. It describes the variability of an estimate rather than the variability of individual observations. Because the underlying population distribution is usually unknown, reported standard errors are generally estimates calculated from sample data. (online.stat.psu.edu)
Definition and interpretation
Let be a statistic used as an estimator of a population parameter . Treating as a random variable, its standard error is
where denotes variance and denotes expected value. The standard error has the same units as the statistic. A smaller value indicates less variation across repeated samples from the specified population and design. (stat.berkeley.edu)
Standard error measures precision, not necessarily accuracy. An estimator may have a small standard error while remaining systematically displaced from its target. For a fixed parameter, mean squared error separates these components:
Thus, estimator bias is not included in the standard error. Sampling precision alone does not establish that an estimate is close to the intended population value. (stat.berkeley.edu)
Standard error of the mean
For observations with statistical independence, a common population mean , and finite variance , the sample mean satisfies
This variance identity does not require a normal distribution. Normality becomes relevant when specifying the exact distribution of the mean or constructing particular inferential procedures. (online.stat.psu.edu)
When is unknown, the conventional estimate is
The denominator incorporates Bessel’s correction. Although is unbiased for the population variance under independent, identically distributed sampling, taking its square root does not generally preserve unbiasedness. The estimated standard error is itself subject to sampling variation. (stat.berkeley.edu)
For illustration, a sample of 100 observations with has an estimated standard error of . With otherwise comparable observations, quadrupling the sample size halves the theoretical standard error. The standard deviation of individual observations need not decrease as the sample grows. (online.stat.psu.edu)
Proportions and regression coefficients
For independent Bernoulli observations with success probability , the sample proportion has
A common plug-in estimate replaces with . This follows from the binomial distribution of the number of successes and the fact that a proportion is a mean of zero–one observations. (online.stat.psu.edu)
In linear regression, each estimated coefficient has its own standard error. For ordinary least squares with fixed, full-rank design matrix , uncorrelated errors, and common error variance ,
Coefficient standard errors are the square roots of the diagonal entries of this covariance matrix, usually after estimating from residuals. They concern coefficient uncertainty, not the spread of residuals or uncertainty in an individual future outcome. (stata.com)
Confidence intervals and hypothesis tests
Standard errors provide the scale for many confidence intervals. A common approximate form is an estimate plus or minus a critical multiplier times its estimated standard error. For a normally distributed population with unknown variance, the exact interval for the mean is
using Student’s t-distribution with degrees of freedom. A confidence level describes the long-run coverage of the interval-producing procedure, not a probability assigned to a fixed parameter after observing the data. (itl.nist.gov)
The central limit theorem supports normal approximations for means and proportions under suitable conditions. In statistical hypothesis testing, a difference between an estimate and a hypothesized value is often divided by a standard error to form a test statistic. For a proportion test, the denominator commonly uses the null-hypothesis probability rather than the observed proportion. (online.stat.psu.edu)
Dependence and alternative estimation methods
Standard errors depend on the sampling design and assumptions about dependence. The independent-observation formula can be inappropriate for clustered observations or time series. Heteroskedasticity-robust, cluster-robust, and autocorrelation-robust estimators accommodate different departures from conventional regression assumptions; none is a universal correction for every dependence structure. (stata.com)
When an analytical formula is unavailable, bootstrap sampling can estimate a standard error by repeatedly resampling the observed data and recomputing the statistic. If the resulting estimates are , their sample standard deviation estimates the statistic’s sampling variability:
This standard deviation is not divided by : the replicates approximate the distribution of the statistic itself, rather than forming a new sample whose mean is the target. (docs.scipy.org)