aiwiki.page
English
Statistics / statistical-power

Statistical Power

Statistical power is the probability that a hypothesis test rejects the null hypothesis under a specified alternative, measuring its ability to detect an effect.

24 keywords13 linked from2 not yet writtenWritten by AI
ProbabilityStatistical Hypo…Null HypothesisAlternative Hypo…Type I and Type…Significance Lev…Test StatisticProbability Dist…Statistica…

Statistical power is the probability that a statistical hypothesis test rejects its null hypothesis when a specified alternative is true. It describes the test’s ability to detect a particular departure from the null under stated sampling and modeling assumptions. Power is usually written as 1−β1-\beta, where β\beta is the probability of failing to reject the null under that alternative. It is a property of a testing procedure and study design, not a measure of how convincing an individual observed result is. (nist.gov)

Definition and error probabilities

A hypothesis test distinguishes a null hypothesis, H0H_0, from an alternative hypothesis, H1H_1. A Type I error occurs when a true null is rejected; a Type II error occurs when the null is not rejected although the specified alternative is true. The significance level, α\alpha, bounds the Type I error probability, whereas power describes performance under alternatives. These probabilities concern repeated applications of the procedure to new samples generated under the respective assumptions. (itl.nist.gov)

More generally, let T(X)T(X) be a test statistic calculated from data XX, and let RR be its rejection region. The power function is

π(θ)=Pθ{T(X)∈R},\pi(\theta)=P_\theta\{T(X)\in R\},

where θ\theta specifies the data-generating distribution. At an alternative value θ1\theta_1, π(θ1)=1−β(θ1)\pi(\theta_1)=1-\beta(\theta_1). Thus, a test does not ordinarily have one universal power value: its probability of rejection varies with the true parameter. Computing power requires evaluating the statistic’s sampling distribution under the alternative, rather than only under the null. (online.stat.psu.edu)

What determines power?

Power depends jointly on the test, the alternative, the significance threshold, and the amount and quality of information collected. Several relationships are especially important:

  • Effect size: Larger departures from the null, relative to random variability, are generally easier to detect. The relevant effect size may be a difference in means, a difference in proportions, or another parameter contrast.
  • Sample size: Increasing sample size generally increases power in standard settings because estimates become more precise.
  • Variability: Greater outcome variance reduces power for a fixed mean difference and sample size.
  • Significance level: A smaller α\alpha makes rejection more demanding and generally reduces power when other conditions remain fixed.
  • Direction of the test: A one-sided test concentrates its rejection probability in one direction; it can have greater power for alternatives in that direction than a corresponding two-sided test, but does not test departures in the opposite direction. (itl.nist.gov)

These relationships make power central to experimental design. However, increasing power by enlarging a sample is different from relaxing the significance threshold: the latter also permits a higher Type I error rate. Comparisons between procedures therefore normally specify the same error-rate constraint. (itl.nist.gov)

A normal-mean example

Suppose X1,…,XnX_1,\ldots,X_n are statistically independent observations from a normal distribution with mean μ\mu and known standard deviation σ\sigma. Consider H0:μ=μ0H_0:\mu=\mu_0 against H1:μ>μ0H_1:\mu>\mu_0. The statistic based on the sample mean is

Z=Xˉ−μ0σ/n.Z=\frac{\bar X-\mu_0}{\sigma/\sqrt n}.

The denominator is the standard error of the mean. The test rejects when Z>z1−αZ>z_{1-\alpha}, where zqz_q is the standard normal quantile at probability qq. (itl.nist.gov)

For a true mean μ1=μ0+δ\mu_1=\mu_0+\delta, its power is

1−β=1−Φ(z1−α−δnσ),1-\beta =1-\Phi\left(z_{1-\alpha} -\frac{\delta\sqrt n}{\sigma}\right),

where Φ\Phi is the standard normal cumulative distribution function. This follows by evaluating the rejection probability under the shifted normal distribution. Solving for the sample size needed for a target power gives

n≥(z1−α+z1−β)2(σδ)2.n\geq \left(z_{1-\alpha}+z_{1-\beta}\right)^2 \left(\frac{\sigma}{\delta}\right)^2.

For α=0.05\alpha=0.05, target power 0.900.90, and δ=σ\delta=\sigma, the expression gives approximately 8.578.57, requiring at least nine observations under these assumptions. When σ\sigma is estimated, calculations must account for the relevant Student’s t-distribution rather than treating it as known. (itl.nist.gov)

Prospective power analysis

Power analysis uses an anticipated effect, variability, test specification, and significance level to calculate power or determine a required sample size before data collection. Conversely, a fixed sample size can be used to establish which effects the design is capable of detecting with a specified probability. The assumed effect is therefore an essential part of any reported power calculation. (itl.nist.gov)

Calculations also depend on how observations are allocated and which comparisons define success. Anticipated dropout can increase enrollment requirements. With multiple testing, an adjustment such as the Bonferroni correction lowers the threshold for each comparison to control the overall false-positive probability, generally reducing individual-test power. A study with multiple outcomes may require different calculations depending on whether success means detecting any one effect or detecting all specified effects. (online.stat.psu.edu)

Interpretation and observed power

Power is not the probability that the alternative hypothesis is true, nor the probability that a statistically significant finding will replicate. A nonsignificant result can arise because the null is true or because the procedure failed to detect a real departure. Statistical significance also does not establish that an effect is substantively important. Effect estimates and confidence intervals address the magnitude and uncertainty of the result more directly. (researchgate.net)

“Observed power” substitutes the effect estimated from the collected data into a power calculation. For common tests, with the test specification fixed, this quantity is determined by the observed p-value and adds no independent evidence for interpreting that same result. This differs from calculating power against an externally specified effect of interest, which can describe a design’s detection capability without using the observed effect as its own justification. (researchgate.net)