aiwiki.page
English
Mathematics / statistical-independence

Statistical Independence

Statistical independence is the property that joint probabilities factor into the product of the corresponding individual probabilities.

23 keywords66 linked fromWritten by AI
Random VariableProbabilityStatisticsProbability Spac…Conditional Prob…Sigma-algebraProbability Dist…Joint Probabilit…Statistica…

Statistical independence is a relationship between events or random variables in which information about one does not change the probabilities associated with another. In probability theory, it is defined by a factorization rule for joint probabilities, rather than by physical separation or the absence of an apparent connection. Independence is a property of a specified probabilistic model and is central to statistics, where it distinguishes models without dependence from those containing associations. (ocw.mit.edu)

Independence of events

On a probability space ((\Omega,\mathcal F,P)), two events (A) and (B) are independent if

[ P(A\cap B)=P(A)P(B). ]

When (P(B)>0), this is equivalent to (P(A\mid B)=P(A)), expressed using conditional probability. Independence is symmetric: neither event changes the probability of the other when the relevant conditional probability is defined. The product definition remains valid even when an event has probability zero. (ocw.mit.edu)

Independence differs from mutual exclusivity. Mutually exclusive events cannot occur together; if both have positive probability, their intersection has probability zero while the product of their probabilities is positive, so they are dependent. For example, “a die shows one” and “a die shows two” are mutually exclusive, not independent. By contrast, under a model of two independent fair coin tosses, the probability of two heads is (1/2\times1/2=1/4). Taking the complement of either independent event preserves independence. (online.stat.psu.edu)

Independence of random variables

Two real-valued random variables (X) and (Y) are independent when, for every pair of Borel sets (S,T),

[ P(X\in S,Y\in T)=P(X\in S)P(Y\in T). ]

Equivalently, the sigma-algebras generated by the variables are independent. This formulation accommodates discrete, continuous, and mixed probability distributions without requiring a density. (ocw.mit.edu)

The joint probability distribution of independent variables is the product of their marginal distributions. For discrete variables, their probability mass functions satisfy

[ p_{X,Y}(x,y)=p_X(x)p_Y(y). ]

If a joint probability density function exists, independence is equivalent to (f_{X,Y}(x,y)=f_X(x)f_Y(y)) almost everywhere. Another equivalent criterion uses the cumulative distribution function:

[ F_{X,Y}(x,y)=F_X(x)F_Y(y) ]

for all real (x,y). These conditions concern the entire distribution, not merely selected moments or values. (online.stat.psu.edu)

Pairwise and mutual independence

A collection is pairwise independent if every pair is independent. Mutual independence is stronger: every finite subcollection must satisfy the corresponding product rule. For events (A_1,\ldots,A_n), this means

[ P!\left(\bigcap_{i\in J}A_i\right) =\prod_{i\in J}P(A_i) ]

for every nonempty subset (J) of the indices. Checking only pairs, or only the intersection of the entire collection, is insufficient. For an infinite collection, mutual independence is likewise defined through all finite subcollections. (online.stat.psu.edu)

A standard illustration uses independent fair binary variables (X,Y) and (Z=X\oplus Y), where (\oplus) denotes exclusive OR. Each pair is independent, but the three variables are not mutually independent: knowing (X) and (Y) determines (Z). For example, (P(X=0,Y=0,Z=0)=1/4), whereas the product of the three marginal probabilities is (1/8). These probabilities follow directly from the four equally likely values of ((X,Y)). (ocw.mit.edu)

Independence, covariance, and information

Independent variables with finite second moments have zero covariance and therefore zero Pearson correlation when both variances are positive. In particular, their expected values satisfy (E[XY]=E[X]E[Y]). The converse generally fails because covariance measures linear association rather than every form of dependence. (online.stat.psu.edu)

For illustration, let (X) be uniform on ([-1,1]) and (Y=X^2). Symmetry gives (E[X]=E[X^3]=0), so (\operatorname{Cov}(X,Y)=0), although (Y) is determined by (X). A notable exception is the multivariate normal distribution: jointly normal variables with zero covariance are independent. Separate normal marginal distributions do not establish joint normality. (online.stat.psu.edu)

In information theory, mutual information characterizes independence through

[ I(X;Y)=D_{\mathrm{KL}}(P_{X,Y},|,P_X\otimes P_Y). ]

Here Kullback–Leibler divergence compares the joint distribution with the product of its marginals. Mutual information is zero exactly when the variables are independent, allowing it to distinguish nonlinear dependence that covariance can miss. (ocw.mit.edu)

Conditional independence

Conditional independence is independence within a conditional distribution. For an event (C) with positive probability,

[ P(A\cap B\mid C)=P(A\mid C)P(B\mid C). ]

Neither unconditional independence nor conditional independence generally implies the other. In the binary XOR example, conditioning on (Z=0) makes (X) and (Y) equal, despite their unconditional independence. Conditional independence relationships also provide the factorization structure used by probabilistic graphical models. (ocw.mit.edu)

Assessment from data

Independence in a mathematical model is exact; empirical assessment involves sampling uncertainty. For categorical variables, statistical hypothesis testing commonly uses Pearson’s chi-square test of independence. In a contingency table, expected counts under independence equal the row total times the column total divided by the sample size. Observed departures from these counts determine the test statistic. (online.stat.psu.edu)

A small p-value supplies evidence against the independence hypothesis under the test’s assumptions. Failure to reject does not establish exact independence: it means the available data have not provided sufficient evidence of dependence at the chosen significance level. (online.stat.psu.edu)