An alternative hypothesis is a proposition in statistical hypothesis testing that specifies departures from a null hypothesis. Usually denoted , , or , it describes parameter values or data-generating distributions that a test is designed to detect. For example, a null hypothesis may state that a population mean equals a specified value, while the alternative states that it differs from that value. The alternative helps determine which observations count as evidence against the null. (itl.nist.gov)
Mathematical formulation
In statistics, a statistical hypothesis concerns a population or its probability distribution, rather than merely describing an observed sample. Suppose a model is indexed by a parameter belonging to a parameter space . A testing problem can be written
where and are disjoint. In an exhaustive formulation, : the alternative contains every possibility excluded by the null. Other testing problems compare specified, disjoint possibilities without requiring them to exhaust the model. (itl.nist.gov)
A simple hypothesis specifies the distribution completely; a composite hypothesis leaves more than one distribution possible. For example, if observations follow a normal distribution with known variance, is simple, whereas is composite. Specifying the mean alone does not make a hypothesis simple if the variance remains unknown. This distinction matters because a test’s performance can vary across the distributions contained in a composite alternative. (online.stat.psu.edu)
One-sided and two-sided alternatives
For a population mean and reference value , common formulations are:
- Two-sided: against .
- Upper-sided: against .
- Lower-sided: against .
A two-sided alternative includes departures in either direction. A one-sided alternative includes departures only in the specified direction. Introductory presentations sometimes write the one-sided null as equality because the boundary value determines the usual critical value; the inequality formulation makes the complementary hypotheses explicit. (itl.nist.gov)
The alternative determines the relevant rejection region. In a conventional symmetric two-sided test, extreme values in both tails count against the null, commonly with half the significance level allocated to each tail. An upper-sided test instead rejects for sufficiently large values of its test statistic. Thus, the same observations may produce different decisions under different alternatives: the tests address different questions. (itl.nist.gov)
Role in constructing tests
A test combines a statistic with a decision rule calibrated using its sampling distribution under the null. The p-value measures how extreme the observed statistic is relative to the null model, with “extreme” defined by the test and its alternative. It is not calculated simply by asking whether a sample estimate differs numerically from the hypothesized value. Sampling variability must also be considered. (itl.nist.gov)
For example, for independent normal observations with unknown population standard deviation, a one-sample mean test uses
where is the sample mean, the sample standard deviation, and the sample size. The denominator is the estimated standard error. At the null boundary, follows Student’s t-distribution with degrees of freedom. The alternative determines whether rejection uses the upper tail, lower tail, or both tails. (itl.nist.gov)
The alternative also enters optimal-test theory. The Neyman–Pearson lemma establishes that, for a simple null against a simple alternative, an appropriate ratio of the two models’ likelihood functions yields a most powerful test at a specified size. For composite alternatives, one test need not be most powerful at every parameter value. (online.stat.psu.edu)
Errors and statistical power
The two principal testing errors are rejecting a true null hypothesis and failing to reject it when the alternative holds. A level- test controls the first error probability at no more than across the null parameter values. For a particular alternative value , statistical power is
The corresponding Type II error probability is . Consequently, a composite alternative generally has a power function, not one universal power value. (itl.nist.gov)
Power depends on the actual effect size, sample size, variability, significance threshold, and testing procedure. In standard mean-comparison settings, larger samples or larger departures from the null generally increase power. Power calculations therefore require a specified alternative effect, often one considered substantively meaningful, rather than the statement that “some difference exists.” (online.stat.psu.edu)
Interpretation and reporting
Rejecting the null provides evidence against it under the test’s assumptions; it does not establish the alternative with certainty. In particular, neither nor is the probability that the alternative is true. Statistical significance also does not measure the magnitude or practical importance of a departure. (amstat.org)
Failure to reject likewise does not demonstrate equality. A matching confidence interval shows which parameter values remain compatible with the data under the procedure. For the conventional two-sided mean test, rejection at level corresponds to exclusion of from the associated interval, subject to matching boundary conventions. (itl.nist.gov)
The inferential procedure must account for multiple testing and selective reporting. Searching among many alternatives and reporting only a favorable result can distort the evidence. Hypotheses, model assumptions, analytical choices, and uncertainty estimates are therefore relevant parts of reporting, not information replaced by a single rejection decision. (magazine.amstat.org)