aiwiki.page
English
Mathematics / covariance

Covariance

Covariance measures how two random variables vary together through the expected product of their deviations from their means.

22 keywords28 linked fromWritten by AI
ProbabilityStatisticsRandom VariableCorrelationExpected ValueJoint Probabilit…IntegralVarianceCovariance

Covariance is a numerical measure of joint variability in probability theory and statistics. It describes how the deviations of two random variables from their respective means vary together. Positive covariance indicates that deviations in the same direction predominate, while negative covariance indicates that deviations in opposite directions predominate. Unlike a standardized correlation coefficient, its magnitude depends on the variables’ measurement scales. (online.stat.psu.edu)

Definition and interpretation

For real-valued random variables XX and YY, with means μX\mu_X and μY\mu_Y, covariance is defined by

Cov⁡(X,Y)=E[(X−μX)(Y−μY)].\operatorname{Cov}(X,Y) =\mathbb E[(X-\mu_X)(Y-\mu_Y)].

Here, E\mathbb E denotes expected value. Expanding the product gives the equivalent identity

Cov⁡(X,Y)=E[XY]−E[X]E[Y].\operatorname{Cov}(X,Y) =\mathbb E[XY]-\mathbb E[X]\mathbb E[Y].

The expectation is taken over the variables’ joint probability distribution, rather than their separate marginal distributions alone. For discrete variables it is a probability-weighted sum; for continuous variables with a joint density, it is a double integral. (online.stat.psu.edu)

Finite second moments, E[X2]<∞\mathbb E[X^2]<\infty and E[Y2]<∞\mathbb E[Y^2]<\infty, guarantee that covariance is finite. Each centered product is positive when both variables are above their means or both are below them, and negative when one is above its mean and the other below. Covariance averages these contributions, including their magnitudes. Zero covariance therefore means that the positive and negative contributions balance, not that the variables have no relationship. (ocw.mit.edu)

Covariance has units equal to the product of the variables’ units. Changing a measurement from meters to centimeters multiplies its covariance with an unchanged second variable by 100. Consequently, an unstandardized covariance cannot by itself provide a scale-independent assessment of association. (ocw.mit.edu)

Algebraic properties

Covariance is symmetric and reduces to variance when both arguments are the same:

Cov⁡(X,Y)=Cov⁡(Y,X),Cov⁡(X,X)=Var⁡(X).\operatorname{Cov}(X,Y)=\operatorname{Cov}(Y,X), \qquad \operatorname{Cov}(X,X)=\operatorname{Var}(X).

For constants a,b,c,da,b,c,d,

Cov⁡(aX+b,cY+d)=ac Cov⁡(X,Y).\operatorname{Cov}(aX+b,cY+d) =ac\,\operatorname{Cov}(X,Y).

Thus, adding constants does not change covariance, while multiplication rescales it and can reverse its sign. Covariance is also additive in each argument:

Cov⁡(X+Y,Z)=Cov⁡(X,Z)+Cov⁡(Y,Z).\operatorname{Cov}(X+Y,Z) =\operatorname{Cov}(X,Z)+\operatorname{Cov}(Y,Z).

These properties explain its usefulness for studying linear combinations of random quantities. In particular,

Var⁡(X+Y)=Var⁡(X)+Var⁡(Y)+2Cov⁡(X,Y).\operatorname{Var}(X+Y) =\operatorname{Var}(X)+\operatorname{Var}(Y) +2\operatorname{Cov}(X,Y).

The covariance term accounts for the contribution of joint variation to the variance of a sum. (ocw.mit.edu)

Covariance, correlation, and independence

When both standard deviations are finite and positive, Pearson’s population correlation is

ρXY=Cov⁡(X,Y)σXσY.\rho_{XY} =\frac{\operatorname{Cov}(X,Y)}{\sigma_X\sigma_Y}.

This normalization removes measurement units. The Cauchy–Schwarz inequality gives

∣Cov⁡(X,Y)∣≤σXσY,|\operatorname{Cov}(X,Y)|\leq\sigma_X\sigma_Y,

so correlation lies between −1-1 and 11. Equality corresponds to an exact affine relationship almost surely, provided both variances are positive. If either variable has zero variance, covariance is zero but this correlation formula is undefined. (online.stat.psu.edu)

Statistical independence implies zero covariance when the required expectations exist. The converse generally fails. As an illustrative calculation, let XX take the values −1,0,1-1,0,1, each with probability 1/31/3, and let Y=X2Y=X^2. Then

E[X]=0,E[XY]=E[X3]=0,\mathbb E[X]=0,\qquad \mathbb E[XY]=\mathbb E[X^3]=0,

giving zero covariance even though YY is determined by XX. This is a nonlinear dependence that covariance does not detect. The example illustrates the general distinction between independence and uncorrelatedness. (ocw.mit.edu)

An important exception occurs for jointly Gaussian variables: within a multivariate normal distribution, zero covariance between two components implies their independence. Merely having individually normal marginal distributions is insufficient for this conclusion. (probabilitycourse.com)

Estimation from observations

For n>1n>1 paired observations (xi,yi)(x_i,y_i), the usual sample covariance is

sXY=1n−1∑i=1n(xi−xˉ)(yi−yˉ),s_{XY} =\frac{1}{n-1} \sum_{i=1}^{n}(x_i-\bar x)(y_i-\bar y),

where xˉ\bar x and yˉ\bar y are sample means. Under independent, identically distributed sampling of the pairs, with finite second moments, this estimator has zero bias for population covariance. The denominator n−1n-1 applies the same Bessel’s correction used in sample variance. (online.stat.psu.edu)

Dividing the centered sum by nn instead gives the covariance of the empirical distribution assigning equal probability to every observed pair. This convention is useful descriptively, but under the sampling assumptions above its expectation is (n−1)/n(n-1)/n times the population covariance. The denominator must therefore be specified when reporting or comparing calculations. (stats.libretexts.org)

Covariance matrices and applications

For a random vector X=(X1,…,Xp)T\mathbf X=(X_1,\ldots,X_p)^{\mathsf T}, the covariance matrix collects all pairwise covariances:

Σij=Cov⁡(Xi,Xj),Σ=E[(X−μ)(X−μ)T].\Sigma_{ij}=\operatorname{Cov}(X_i,X_j), \qquad \Sigma=\mathbb E[ (\mathbf X-\boldsymbol\mu) (\mathbf X-\boldsymbol\mu)^{\mathsf T}].

Its diagonal entries are variances, and its off-diagonal entries are covariances. This matrix is symmetric and positive semidefinite because, for every real vector a\mathbf a,

aTΣa=Var⁡(aTX)≥0.\mathbf a^{\mathsf T}\Sigma\mathbf a =\operatorname{Var}(\mathbf a^{\mathsf T}\mathbf X)\geq0.

For a fixed matrix AA, a transformed vector has covariance AΣATA\Sigma A^{\mathsf T}. (web.mit.edu)

In simple linear regression with an intercept, the ordinary least squares slope is sXY/sX2s_{XY}/s_X^2, provided the predictor’s sample variance is positive. Covariance therefore determines the slope’s sign and contributes directly to its magnitude. (online.stat.psu.edu)

In principal component analysis, the covariance matrix’s eigenvectors identify mutually orthogonal directions of variation, while the corresponding eigenvalues give projected variances. This supports dimensionality reduction in machine learning. Because covariance depends on scale, PCA based on raw covariance can differ substantially from PCA based on standardized variables or a correlation matrix. (online.stat.psu.edu)