A test statistic is a quantity calculated from sample data for use in statistical hypothesis testing. It summarizes an aspect of the observations relevant to distinguishing a null hypothesis from an alternative hypothesis. Its observed value is compared with a reference distribution describing its behavior under the null hypothesis and associated assumptions. This comparison determines whether the result falls within a rejection region or yields a small p-value. (online.stat.psu.edu)
Mathematical definition
For observations , a test statistic can be written as
It is a specified function of the sample, not a function of unknown parameter values. Hypothesized parameter values and quantities estimated from the observations may nevertheless appear in its formula. Before the observations are collected, is a random variable; after collection, its realized value is denoted . Common statistics measure a standardized difference, a discrepancy between observed and expected counts, or a comparison of fitted models. (stat.berkeley.edu)
The sampling distribution of under the null hypothesis is called its null distribution. A number alone does not establish how unusual a result is: its interpretation depends on this distribution and on which values count as evidence against the null. Large values, small values, or values in both tails may be relevant, depending on the alternative hypothesis. (stat.berkeley.edu)
Rejection regions and p-values
A rejection region specifies the statistic values that lead to rejection. A test with significance level controls the probability of rejecting a true null hypothesis at no more than , under its assumptions. For a simple null hypothesis, this is expressed as
Boundaries of the region are called critical values. This control concerns the first of the Type I and Type II errors; it does not guarantee a correct decision for any individual sample. (itl.nist.gov)
For an upper-tailed test with a specified null distribution,
For a two-sided test based on a symmetric, zero-centered statistic, a common definition is
Thus, the statistic and p-value are distinct: the former summarizes the data, while the latter evaluates its extremeness using a probability distribution. Two-sided extremeness depends on the testing procedure rather than on a universal rule. (itl.nist.gov)
Standardized statistics for means
Many familiar statistics have the structure
The numerator measures departure from the null value; the denominator expresses sampling variability. For a population mean with known standard deviation , the statistic is
where is the sample mean and is the hypothesized mean. With independent observations from a normal distribution, has a standard normal null distribution. (itl.nist.gov)
When the population standard deviation is unknown, the one-sample t-statistic replaces with the sample standard deviation :
Here, estimates the standard error of the mean. For independent, identically distributed normal observations, the null distribution is Student’s t-distribution with degrees of freedom. Estimating the variability therefore changes the appropriate reference distribution. (itl.nist.gov)
As a numerical illustration, let , , , and . Then : the sample mean lies two estimated standard errors above the hypothesized value. Whether this leads to rejection depends on the chosen significance level, direction of the alternative, and t-distribution with 24 degrees of freedom. The calculation itself is not a decision rule. (itl.nist.gov)
Discrepancy and likelihood statistics
For categorical counts, Pearson’s chi-squared statistic is
where and are observed and null-model expected counts. Larger values indicate greater discrepancy. In a goodness-of-fit test, its reference chi-squared distribution is generally an approximation; the degrees of freedom depend on the categories and parameters estimated from the data. Adequate expected counts are important for this approximation. (itl.nist.gov)
Another approach uses the likelihood function. A likelihood-ratio test compares the best fit permitted by the null hypothesis with the best fit in a larger model:
Small values of , or large values of , indicate that imposing the null substantially reduces the achievable likelihood. This approach connects testing with maximum likelihood estimation. (online.stat.psu.edu)
Calibration and interpretation
A null distribution need not come from a standard analytic formula. A permutation test recalculates a chosen statistic over rearrangements justified by the null hypothesis. Enumeration or Monte Carlo sampling can provide its reference distribution. The validity of this calibration depends on the permissible rearrangements, not simply on the ability to shuffle observations. (stat.berkeley.edu)
A statistic’s interpretation remains conditional on the model and sampling assumptions. An extreme result may reflect an incorrect null hypothesis or incorrect assumptions. Moreover, a small p-value does not give the probability that the null hypothesis is true and does not measure effect size or practical importance. Failure to reject is likewise not proof of the null. Estimates and confidence intervals address questions about magnitude and uncertainty that a test decision alone does not resolve. (stat.berkeley.edu)