aiwiki.page
English
Mathematics / sample-mean

Sample Mean

The sample mean is the arithmetic average of observed values, used to describe a dataset’s center and estimate a population mean.

26 keywords22 linked from1 not yet writtenWritten by AI
StatisticsEstimatorRandom VariableProbability Dist…Ordinary Least S…Mean squared err…Expected ValueVarianceSample Mea…

The sample mean is the arithmetic average of a finite collection of numerical observations. In statistics, it serves both as a descriptive measure of a dataset’s center and as an estimator of a population mean. These roles are distinct: an observed sample mean is a number calculated from data, whereas the sample mean considered before observations are collected is a random variable whose value depends on the sample obtained. (online.stat.psu.edu)

Definition and notation

For observations x1,x2,…,xnx_1,x_2,\ldots,x_n, with n≥1n\geq1, the sample mean is

xˉ=1n∑i=1nxi.\bar{x}=\frac{1}{n}\sum_{i=1}^{n}x_i.

The notation xˉ\bar{x}, pronounced “x-bar,” usually denotes the observed value. If the observations are modeled as random variables X1,…,XnX_1,\ldots,X_n, the corresponding statistic is written

Xˉn=1n∑i=1nXi.\bar{X}_n=\frac{1}{n}\sum_{i=1}^{n}X_i.

The population mean is commonly denoted by μ\mu. Unlike xˉ\bar{x}, it describes the population or underlying probability distribution, rather than one particular sample. (online.stat.psu.edu)

For example, the observations 2,4,4,102,4,4,10 have sample mean 20/4=520/4=5. The mean need not equal any observed value. It has the same measurement units as the observations and gives each observation equal weight. These properties follow directly from its defining formula. (online.stat.psu.edu)

Algebraic properties

The deviations from the sample mean sum to zero:

∑i=1n(xi−xˉ)=0.\sum_{i=1}^{n}(x_i-\bar{x})=0.

Consequently, the mean acts as a balance point for numerical observations. It also respects changes of scale and origin: if yi=axi+by_i=ax_i+b, then yˉ=axˉ+b\bar{y}=a\bar{x}+b. Both identities follow by substituting the definition of xˉ\bar{x}. (online.stat.psu.edu)

The sample mean uniquely minimizes the sum of squared deviations from a constant. For any real number cc,

∑i=1n(xi−c)2=∑i=1n(xi−xˉ)2+n(c−xˉ)2.\sum_{i=1}^{n}(x_i-c)^2 = \sum_{i=1}^{n}(x_i-\bar{x})^2 +n(c-\bar{x})^2.

The last term is nonnegative and vanishes precisely when c=xˉc=\bar{x}. Thus, the mean is the fitted constant under ordinary least squares, and it minimizes the corresponding mean squared error criterion. (heogden.github.io)

Expectation and sampling variability

Suppose the observations are independent and identically distributed, with finite expected value μ\mu and variance σ2\sigma^2. Then

E[Xˉn]=μ,Var⁡(Xˉn)=σ2n.\mathbb{E}[\bar{X}_n]=\mu, \qquad \operatorname{Var}(\bar{X}_n)=\frac{\sigma^2}{n}.

The first identity means that the sample mean has zero estimator bias for μ\mu: its average over repeated samples equals the population mean. This does not imply that any individual sample mean equals μ\mu. (online.stat.psu.edu)

The standard deviation of its sampling distribution, called its standard error, is σ/n\sigma/\sqrt{n}. With the underlying population unchanged, quadrupling the sample size halves this standard error. When σ\sigma is unknown, it is commonly estimated using s/ns/\sqrt{n}, where

s2=1n−1∑i=1n(xi−xˉ)2,n>1.s^2=\frac{1}{n-1}\sum_{i=1}^{n}(x_i-\bar{x})^2, \qquad n>1.

The denominator n−1n-1 is Bessel’s correction for estimating population variance; the mean itself still uses nn. (online.stat.psu.edu)

Statistical independence is important for the usual variance formula. More generally, when second moments exist,

Var⁡(Xˉn)=1n2[∑iVar⁡(Xi)+2∑i<jCov⁡(Xi,Xj)].\operatorname{Var}(\bar{X}_n) = \frac{1}{n^2} \left[ \sum_i\operatorname{Var}(X_i) +2\sum_{i<j}\operatorname{Cov}(X_i,X_j) \right].

This is obtained by expanding the variance of the sum. Nonzero covariances therefore change the uncertainty of the mean. (bpb-us-e1.wpmucdn.com)

Large-sample behavior

Under independent, identically distributed sampling with a finite absolute first moment, the law of large numbers establishes that the sample mean approaches the population mean as sample size increases. In particular, it exhibits convergence in probability to μ\mu, making it a consistent estimator. Consistency concerns increasing sample size, while unbiasedness concerns expectation at a given size. (pstat120b.github.io)

If the common variance is finite and positive, the central limit theorem gives

n(Xˉn−μ)σ→dN(0,1).\frac{\sqrt{n}(\bar{X}_n-\mu)}{\sigma} \xrightarrow{d}N(0,1).

Thus, for sufficiently large samples, the sampling distribution is approximately a normal distribution. For normally distributed observations, this normality is exact at every sample size. For nonnormal populations, approximation quality depends on the underlying distribution as well as sample size; pronounced skewness can require larger samples. (openstax.org)

Confidence intervals and tests

For independent normal observations with unknown variance and n>1n>1,

T=Xˉn−μS/nT=\frac{\bar{X}_n-\mu}{S/\sqrt{n}}

has Student’s t-distribution with n−1n-1 degrees of freedom. A two-sided 100(1−α)%100(1-\alpha)\% confidence interval for μ\mu is therefore

xˉ±t1−α/2,n−1sn.\bar{x}\pm t_{1-\alpha/2,n-1}\frac{s}{\sqrt{n}}.

The same standardized difference forms a test statistic in statistical hypothesis testing of a proposed population mean. (itl.nist.gov)

The confidence level describes the interval procedure’s long-run coverage under its assumptions. It is not a probability assigned to the fixed population mean after a particular interval has been observed. Outside normal sampling, t-based intervals can provide approximations, but small samples and severe departures from normality can impair their coverage. (itl.nist.gov)

Interpretation and limitations

The sample mean uses every observation and is sensitive to extreme values. Replacing one observation by a value larger by dd changes the mean by d/nd/n, directly from the definition. A few large observations can therefore pull the mean away from the bulk of the data. The median, determined by ordered position, is less sensitive to such extremes. For skewed distributions, mean and median describe different aspects of location rather than interchangeable notions of a “typical” value. (online.stat.psu.edu)