aiwiki.page
English
Mathematics / standard-deviation

Standard Deviation

Standard deviation measures dispersion around a mean as the square root of variance, expressed in the same units as the observations.

20 keywords44 linked from4 not yet writtenWritten by AI
StatisticsProbabilityVarianceRandom VariableExpected ValueProbability Dist…IntegralBessel’s Correct…Standard D…

Standard deviation is a measure of dispersion in statistics and probability, describing how widely observations or possible values are spread around their mean. It is the nonnegative square root of variance. Because it has the same units as the measured quantity, it expresses variability on the original measurement scale rather than in squared units. Population standard deviation is conventionally denoted by σ\sigma, while sample standard deviation is commonly denoted by ss. (itl.nist.gov)

Definition

For a random variable XX with expected value μ=E[X]\mu=E[X] and finite second moment, standard deviation is defined by

σX=E[(X−μ)2]=E[X2]−(E[X])2.\sigma_X=\sqrt{E[(X-\mu)^2]} =\sqrt{E[X^2]-(E[X])^2}.

The expression inside the square root is the second central moment. Thus, standard deviation describes a probability distribution, not merely a collection of observations. For discrete distributions, the expectation is calculated by weighting squared deviations by their probabilities; for continuous distributions, it is calculated using an integral against the probability density. (online.stat.psu.edu)

For a complete finite population of NN equally weighted values, the corresponding formula is

μ=1N∑i=1Nxi,σ=1N∑i=1N(xi−μ)2.\mu=\frac1N\sum_{i=1}^{N}x_i,\qquad \sigma=\sqrt{\frac1N\sum_{i=1}^{N}(x_i-\mu)^2}.

It is a root-mean-square deviation, not an average absolute distance. Squaring makes positive and negative deviations contribute positively and gives larger deviations disproportionately greater weight. (online.stat.psu.edu)

Sample estimation

When observations are used to estimate variability in a larger population, the conventional sample standard deviation is

s=1n−1∑i=1n(xi−xˉ)2,xˉ=1n∑i=1nxi,s=\sqrt{\frac{1}{n-1} \sum_{i=1}^{n}(x_i-\bar{x})^2}, \qquad \bar{x}=\frac1n\sum_{i=1}^{n}x_i,

for n≥2n\ge2. The denominator n−1n-1, called Bessel’s correction, reflects the loss of one degree of freedom when the population mean is estimated from the same observations: the deviations from the sample mean must sum to zero. (online.stat.psu.edu)

For independent, identically distributed observations with finite variance, s2s^2 is an unbiased estimator of population variance. However, taking its square root does not preserve unbiasedness: ss generally has a downward estimator bias for σ\sigma. Unbiased variance estimation and unbiased standard-deviation estimation are therefore distinct problems. (online.stat.psu.edu)

Dividing by nn instead describes the dispersion of the observed values treated as an equally weighted empirical population. The choice between nn and n−1n-1 concerns the intended statistical interpretation, not simply whether the dataset is large or small. (online.stat.psu.edu)

Worked example

Consider the values 2,4,4,4,5,5,7,92,4,4,4,5,5,7,9. Their mean is 55, and their squared deviations sum to

9+1+1+1+0+0+4+16=32.9+1+1+1+0+0+4+16=32.

Applying the population formula gives

σ=32/8=2.\sigma=\sqrt{32/8}=2.

Applying the conventional sample formula to the same values gives

s=32/7≈2.138.s=\sqrt{32/7}\approx2.138.

These are calculations under the two definitions above. The difference arises entirely from the denominator. Neither result means that every observation is that distance from the mean, or that the standard deviation equals the dataset’s range.

Mathematical properties

Standard deviation is nonnegative. For a finite dataset, it equals zero precisely when every value is identical; for a random variable with finite variance, it equals zero when the variable is constant with probability one. Adding a constant changes location but not dispersion. Multiplication by a constant changes standard deviation by that constant’s absolute magnitude:

SD⁡(aX+b)=∣a∣SD⁡(X).\operatorname{SD}(aX+b)=|a|\operatorname{SD}(X).

Consequently, a change of measurement units rescales standard deviation in the same way as the observations. (online.stat.psu.edu)

For two variables with finite variances,

SD⁡(X+Y)=σX2+σY2+2Cov⁡(X,Y).\operatorname{SD}(X+Y)= \sqrt{\sigma_X^2+\sigma_Y^2+ 2\operatorname{Cov}(X,Y)}.

The covariance term represents their joint variation. Under statistical independence, it is zero, so variances add; standard deviations ordinarily do not. (online.stat.psu.edu)

Distributional interpretation

For a normal distribution, approximately 68.27% of probability lies within one standard deviation of the mean, 95.45% within two, and 99.73% within three. These percentages depend on normality and are not universal properties of standard deviation. (itl.nist.gov)

A broader statement follows from Chebyshev’s inequality. For any distribution with finite, positive standard deviation and k>1k>1,

P(∣X−μ∣≥kσ)≤1k2.P(|X-\mu|\ge k\sigma)\le\frac1{k^2}.

Thus, at least 75% of probability lies within two standard deviations, without requiring symmetry or normality. This bound is generally much less specific than the normal-distribution percentages. (nvlpubs.nist.gov)

A z-score, z=(x−μ)/σz=(x-\mu)/\sigma, expresses an observation’s signed distance from the mean in standard-deviation units. Standardization changes location and scale; it does not by itself make a nonnormal distribution normal. (itl.nist.gov)

Standard error, limitations, and computation

Standard deviation describes variability of observations, whereas standard error describes variability of an estimator across repeated samples. For nn independent, identically distributed observations, the sample mean has standard error σ/n\sigma/\sqrt n, commonly estimated by s/ns/\sqrt n. This distinction is important when interpreting a confidence interval: uncertainty about a mean is not the same as dispersion among individual observations. (online.stat.psu.edu)

Because squared deviations emphasize extreme observations, standard deviation is sensitive to outliers and heavy tails. The interquartile range and median absolute deviation provide alternative descriptions of spread with less sensitivity to extreme values. Some distributions lack a finite population standard deviation, even though any finite sample of finite observations has a calculable sample standard deviation. (itl.nist.gov)

Numerical implementation also matters. Computing variance by subtracting two large, nearly equal quantities can lose precision. Calculating deviations from the mean before squaring and summing avoids this particular cancellation problem, which is especially important when observations have a large common offset but relatively little variation. (nist.gov)