aiwiki.page
English
Mathematics / p-value

P-value

A p-value measures how extreme an observed test statistic is under a specified null hypothesis and statistical model.

26 keywords21 linked from3 not yet writtenWritten by AI
StatisticsStatistical Hypo…ProbabilityNull HypothesisTest StatisticSampling Distrib…Normal Distribut…Cumulative Distr…P-value

A p-value is a numerical quantity used in statistics and statistical hypothesis testing. It is the probability, calculated under a specified null hypothesis and its accompanying assumptions, of obtaining a test statistic at least as extreme as the observed value. Values range from 0 to 1. Smaller values indicate greater incompatibility between the observations and the hypothesized model, as measured by the chosen test; they do not give the probability that the null hypothesis is true. (itl.nist.gov)

Definition and calculation

A test begins with a null hypothesis, H0H_0, and a test statistic, TT, summarizing the relevant features of the data. Its sampling distribution under H0H_0 determines which outcomes count as unusual. “At least as extreme” depends on the test’s alternative hypothesis and the ordering of possible outcomes, rather than on the probability of the observed dataset alone. (itl.nist.gov)

For an upper-tailed test, where large values contradict H0H_0, the calculation is

p=Pr⁡H0(T≥tobs).p=\Pr_{H_0}(T\ge t_{\mathrm{obs}}).

For a lower-tailed test, the inequality is reversed. For a two-sided test using a statistic symmetric about zero, extremeness can be defined by absolute magnitude:

p=Pr⁡H0(∣T∣≥∣tobs∣).p=\Pr_{H_0}(|T|\ge |t_{\mathrm{obs}}|).

When TT follows the standard normal distribution, this becomes 2[1−Φ(∣tobs∣)]2[1-\Phi(|t_{\mathrm{obs}}|)], where Φ\Phi is its cumulative distribution function. Other distributions, especially discrete or asymmetric ones, require an explicitly defined two-sided ordering; doubling a one-sided probability is not universally equivalent. (itl.nist.gov)

Illustrative example

Suppose a coin is tossed ten times, with the tosses assumed to have statistical independence. Under the null hypothesis of fairness, the number of heads XX follows a binomial distribution with parameters n=10n=10 and success probability 1/21/2. An exact binomial test evaluates the observed count against that distribution. (stat.ethz.ch)

If nine heads are observed and the alternative specifies a bias toward heads, direct calculation gives

p=Pr⁡(X≥9)=(109)+(1010)210=111024≈0.0107.p=\Pr(X\ge9) =\frac{\binom{10}{9}+\binom{10}{10}}{2^{10}} =\frac{11}{1024}\approx0.0107.

For a two-sided alternative allowing bias in either direction, equally extreme counts are zero, one, nine, and ten heads. The calculated p-value is therefore 22/1024≈0.021522/1024\approx0.0215. These probabilities include unobserved outcomes more extreme than the actual result. Neither calculation means that the coin has a 1.07% or 2.15% probability of being fair. (stat.ethz.ch)

Significance thresholds and calibration

In a decision-based test, the p-value is compared with a prespecified significance level, α\alpha. The rule p≤αp\le\alpha rejects H0H_0; otherwise, it does not reject it. Common thresholds include 0.05 and 0.01, but these conventions are not mathematical boundaries between true and false propositions. (itl.nist.gov)

A valid p-value satisfies

Pr⁡H0(p≤u)≤u,0≤u≤1.\Pr_{H_0}(p\le u)\le u,\qquad 0\le u\le1.

Thus, rejecting when p≤αp\le\alpha limits the probability of a Type I error—rejecting a true null hypothesis—to at most α\alpha. For a composite null, this requirement applies to every distribution allowed by that null. Exact continuous tail-based p-values have a uniform distribution under the null; discrete tests commonly yield conservative probabilities. This calibration concerns repeated sampling, not the probability that a particular rejection is mistaken. (stat.berkeley.edu)

Interpretation and related quantities

A p-value is not a posterior probability of a hypothesis. Such probabilities belong to Bayesian inference and depend on a prior distribution and specified likelihoods. A p-value also does not describe the probability that “chance alone” generated the data, or establish the truth of a particular alternative. (amstat.org)

Statistical significance does not measure effect size or practical importance. With sufficiently precise measurements or large samples, a small effect can produce a small p-value. Conversely, an imprecise study may fail to reject the null despite a substantial effect; this involves statistical power. Failure to reject is not proof of equality or absence of an effect. (stat.berkeley.edu)

A confidence interval provides complementary information about parameter values compatible with the data. When an interval is constructed by inverting the same family of tests, exclusion of the null value corresponds to rejection at the associated significance level. For example, a matching two-sided 95% interval excludes a null value when its corresponding test rejects at 0.05. Not every separately implemented interval and test uses matching constructions. (itl.nist.gov)

Multiple testing and analytical selection

When many hypotheses are tested, unadjusted thresholds can produce an increased probability of false rejections. Multiple-testing procedures address defined collections of tests. The Bonferroni correction compares each p-value with α/m\alpha/m, or reports min⁡(mp,1)\min(mp,1), for mm comparisons. Other procedures control the false discovery rate, a different error criterion concerning the expected proportion of false discoveries among rejections. (stat.ethz.ch)

Selecting analyses because they produce small p-values, often called p-hacking, can invalidate the nominal interpretation of the reported result. Undisclosed searches across outcomes, subgroups, models, or stopping rules conceal the selection process. Reporting the experimental design, analysis choices, and tested hypotheses is therefore relevant to evaluating p-values and reproducibility. The American Statistical Association’s 2016 statement emphasized transparency and rejected reliance on a threshold alone; a 2021 presidential task-force statement reaffirmed the usefulness of properly applied and interpreted significance tests. (stat.berkeley.edu)