Type I and Type II errors are the two possible incorrect decisions in statistical hypothesis testing. A Type I error occurs when a true null hypothesis is rejected; a Type II error occurs when a false null hypothesis is not rejected. These categories distinguish false alarms from missed detections and provide a framework in statistics for evaluating testing procedures under repeated sampling. Their probabilities describe a procedure’s behavior under specified conditions, not whether a particular observed conclusion is wrong. (itl.nist.gov)
Decisions and underlying truth
A test contrasts a null hypothesis, , with an alternative hypothesis, . Data are reduced to a test statistic, and a decision rule specifies which values lead to rejection. The underlying hypothesis may be true or false independently of the decision, producing four possibilities: (itl.nist.gov)
| Underlying situation | Reject | Do not reject |
|---|---|---|
| is true | Type I error | Correct non-rejection |
| is false | Correct rejection | Type II error |
“Do not reject” is generally preferable to “accept”: insufficient evidence against a hypothesis does not establish its truth. For example, if states that a manufacturing process has its target mean, rejecting it when the mean is on target is a Type I error; failing to detect an actual departure is a Type II error. The error labels depend on how the null hypothesis is formulated. (itl.nist.gov)
Error probabilities and statistical power
For a simple null hypothesis, the Type I error probability is commonly written
This is a conditional probability evaluated under the null model. A test’s significance level is a specified upper bound on this probability. For a composite null hypothesis, containing several possible parameter values, a level- test must satisfy that bound at every null value. Its actual maximum error probability may be below the nominal level. (stat210a.berkeley.edu)
At a particular alternative parameter value , the Type II error probability is
The corresponding statistical power is
Power is therefore the probability of detecting a specified departure from the null. There is usually no single Type II error probability for an entire composite alternative: different true effect sizes give different detection probabilities. A power curve records this dependence across parameter values. (itl.nist.gov)
Neither nor is generally the probability that a hypothesis is true after observing the data. That reverses the conditioning. Probabilities assigned to hypotheses belong to a different framework, such as Bayesian inference, involving additional model specifications. (amstat.org)
Trade-offs and study design
With the data model, sample size, and test statistic fixed, making rejection harder typically reduces Type I errors while increasing Type II errors. Lowering a significance threshold from 0.05 to 0.01, for example, shrinks the rejection region and reduces power against the same alternative. The two error probabilities are not complements: they are calculated under different underlying situations. (itl.nist.gov)
Increasing sample size can reduce Type II error while maintaining a fixed significance level. Power also depends on variability, effect magnitude, and whether the test is one-sided or two-sided. In experimental design, sample-size calculations specify an alternative of interest together with an acceptable Type I error bound and desired power. Larger effects and lower variance generally make detection easier. (online.stat.psu.edu)
A normal-mean example
Suppose are independent observations from a normal distribution with mean and known standard deviation . Consider against . Using the sample mean , define
The denominator is the standard error of the mean. Under , the statistic’s sampling distribution is standard normal, so rejecting when gives Type I error probability . (online.stat.psu.edu)
At a true mean ,
where is the standard normal cumulative distribution function. This formula shows how increasing , increasing the mean difference, or decreasing reduces missed detections. With and a mean difference equal to three standard errors, the calculated Type II error probability is approximately 0.088, giving power approximately 0.912. (online.stat.psu.edu)
Interpretation and multiple testing
A p-value describes how incompatible the observed statistic is with a specified null model. It is not the probability that the result is a Type I error. Likewise, non-significance does not establish absence of an effect. A confidence interval can indicate which effect magnitudes remain compatible with the data; for matching procedures, exclusion of a hypothesized value corresponds to rejection by the associated test. Statistical significance alone does not measure practical importance. (amstat.org)
In multiple testing, controlling each test separately does not necessarily control the probability of at least one false rejection across the collection. The Bonferroni correction addresses this family-wise error probability by testing each of hypotheses at level . This stricter threshold can reduce power for individual alternatives, illustrating the same false-alarm–missed-detection trade-off at a broader scale. (online.stat.psu.edu)