A sampling distribution is the probability distribution of a statistic calculated from random samples drawn under a specified sampling process. It describes how quantities such as sample means, proportions, or estimated coefficients vary from sample to sample. In statistics, sampling distributions connect observed data with population characteristics and provide the basis for quantifying uncertainty in estimation and testing. They concern the behavior of a statistic across possible samples, rather than the distribution of individual observations within one sample. (online.stat.psu.edu)
Definition and interpretation
Let (X_1,\ldots,X_n) represent sampled observations, and let [ T=g(X_1,\ldots,X_n) ] be a statistic: a function of the sample that does not involve unknown parameters. Because its inputs are random, (T) is a random variable. Its sampling distribution assigns probabilities to its possible values. When (T) is used to estimate a population parameter, it is an estimator; the value calculated from a particular observed sample is an estimate. (stat.berkeley.edu)
Three distributions must be distinguished. The population distribution describes individual values in the population. The empirical distribution describes values in an observed sample. The sampling distribution describes a statistic across possible samples. Repeated sampling is a conceptual device: the distribution exists mathematically even when only one sample is collected. A histogram of statistics from many simulated samples approximates that distribution rather than defining it. (online.stat.psu.edu)
The sampling distribution depends on the population model, sample size, statistic, and sampling design. Sampling with replacement, sampling without replacement, and more complex selection procedures can produce different distributions even for the same statistic. (stat.berkeley.edu)
Center, variability, and estimation
The center of an estimator’s sampling distribution is commonly measured by its expected value. For an estimator (\hat\theta) of a parameter (\theta), its bias is [ \operatorname{Bias}(\hat\theta)=E(\hat\theta)-\theta. ] An unbiased estimator has a sampling distribution whose expectation equals the target parameter; this does not mean every individual estimate equals that parameter. (stat.berkeley.edu)
Its variance measures dispersion across repeated samples. The standard deviation of the sampling distribution is called the standard error: [ \operatorname{SE}(\hat\theta) =\sqrt{\operatorname{Var}(\hat\theta)}. ] A standard error describes uncertainty in a statistic, whereas the standard deviation of individual observations describes variability in the underlying data. In practice, a standard error often must itself be estimated from the sample. (online.stat.psu.edu)
Sampling distribution of the mean
Suppose (X_1,\ldots,X_n) are identically distributed observations satisfying statistical independence, with mean (\mu) and finite variance (\sigma^2). Their sample mean is [ \bar X=\frac{1}{n}\sum_{i=1}^{n}X_i. ] Its expectation and variance are [ E(\bar X)=\mu,\qquad \operatorname{Var}(\bar X)=\frac{\sigma^2}{n}, ] so its standard error is (\sigma/\sqrt n). Consequently, quadrupling the sample size halves the standard error under these assumptions. If the population follows a normal distribution, then [ \bar X\sim N(\mu,\sigma^2/n) ] exactly, for every positive sample size. (online.stat.psu.edu)
For simple random sampling without replacement from a finite population of size (N), the observations are not independent. Defining the population variance using denominator (N), the variance of the sample mean becomes [ \operatorname{Var}(\bar X) =\frac{\sigma^2}{n}\frac{N-n}{N-1}. ] The additional factor accounts for the reduction in uncertainty as the sample approaches the entire population; a census has no sampling variability. (stat.berkeley.edu)
Exact and approximate distributions
The central limit theorem states that, for independent, identically distributed observations with finite positive variance, [ \frac{\sqrt n(\bar X-\mu)}{\sigma} \xrightarrow{d}N(0,1). ] Thus, a normal approximation for the sample mean becomes appropriate as (n) grows. The theorem concerns the standardized mean, not the raw observations. No universal sample-size threshold guarantees an accurate approximation: strongly skewed populations may require substantially larger samples. (online.stat.psu.edu)
For independent binary observations following a Bernoulli distribution with success probability (p), the success count (K) follows a binomial distribution. The sample proportion (\hat p=K/n) therefore has an exact discrete sampling distribution, with [ E(\hat p)=p,\qquad \operatorname{SE}(\hat p)=\sqrt{\frac{p(1-p)}{n}}. ] A normal approximation is useful when both expected successes and expected failures are sufficiently numerous. (online.stat.psu.edu)
For an independent normal sample with unknown variance, the standardized statistic [ T=\frac{\bar X-\mu}{S/\sqrt n} ] has Student’s \(t\)-distribution with (n-1) degrees of freedom, where (S) is the usual sample standard deviation. It is this standardized statistic—not the unstandardized sample mean—that has the (t)-distribution. (stat.umn.edu)
Role in statistical inference
A confidence interval uses sampling-distribution probabilities to construct an interval procedure with specified repeated-sampling coverage. For example, a 95% procedure covers the fixed parameter in 95% of repetitions under its assumptions; the probability statement concerns the procedure before the data are observed. (online.stat.psu.edu)
In statistical hypothesis testing, a test statistic is compared with its sampling distribution under a null hypothesis. A p-value is the probability, under that hypothesis and the assumed model, of a result at least as extreme as the observed result, according to the test’s specified ordering. (online.stat.psu.edu)
Computational approximation
When analytical formulas are unavailable, simulated samples can approximate a sampling distribution. Bootstrap sampling instead uses the observed sample to approximate the unknown population, repeatedly resampling and recalculating the statistic. The resulting bootstrap distribution estimates sampling variability conditional on the observed data; it is not identical to the true sampling distribution. (stat135.berkeley.edu)
The resampling scheme must reflect the original sampling process. Ordinary resampling of individual observations may fail when the sample arose through a more complex design. Increasing the number of bootstrap repetitions improves numerical precision but does not remove deficiencies in the underlying sample or resampling assumptions. (stat20.berkeley.edu)