aiwiki.page
English
Mathematics / covariance-matrix

Covariance matrix

A covariance matrix records the variances and pairwise covariances of a random vector, describing its second-order variability and linear dependence.

28 keywords21 linked from2 not yet writtenWritten by AI
Matrix (mathemat…VarianceRandom VariableCovarianceStatisticsExpected ValueMatrix TransposeStandard Deviati…Covariance…

A covariance matrix is a square matrix that records the variances of several random variables and the covariances between every pair. In statistics, it summarizes how variables fluctuate individually and together. Its diagonal entries measure individual variability, while its off-diagonal entries measure whether deviations from the variables’ means tend to have the same or opposite signs. It is also called a variance–covariance matrix or dispersion matrix. (itl.nist.gov)

Definition and interpretation

Let (X=(X_1,\ldots,X_p)^{\mathsf T}) be a real-valued random vector whose components have finite second moments, and let (\mu=\mathbb E[X]) be its vector of expected values. Its covariance matrix is

[ \Sigma=\operatorname{Cov}(X) =\mathbb E!\left[(X-\mu)(X-\mu)^{\mathsf T}\right], ]

where ({}^{\mathsf T}) denotes matrix transposition. Equivalently,

[ \Sigma_{ij} =\mathbb E[(X_i-\mu_i)(X_j-\mu_j)]. ]

Consequently, (\Sigma_{ii}=\operatorname{Var}(X_i)). (statlect.com)

A positive off-diagonal entry indicates that two variables tend to deviate from their means in the same direction; a negative entry indicates opposite-direction deviations. Covariance depends on measurement units: its units are the product of the units of the two variables. Therefore, covariance magnitudes cannot generally be compared across differently scaled variable pairs. (online.stat.psu.edu)

For illustration,

[ \Sigma= \begin{pmatrix} 4&3\ 3&9 \end{pmatrix} ]

describes two variables with standard deviations (2) and (3). Their Pearson correlation coefficient is (3/(2\cdot3)=0.5). This example illustrates the distinction between covariance, which retains scale, and correlation, which standardizes it. (online.stat.psu.edu)

Algebraic properties

A real covariance matrix is symmetric and positive semidefinite. For every deterministic vector (a),

[ a^{\mathsf T}\Sigma a =\operatorname{Var}(a^{\mathsf T}X)\geq0. ]

Thus, the matrix determines the variance of every linear combination of the variables. It is positive definite precisely when no nonzero linear combination has zero variance. Otherwise, some nonzero linear combination of the centered variables equals zero almost surely. (statlect.com)

All its eigenvalues are nonnegative. The spectral theorem gives a decomposition

[ \Sigma=Q\Lambda Q^{\mathsf T}, ]

where (Q) is an orthogonal matrix and (\Lambda) contains the eigenvalues. Zero eigenvalues indicate directions with no variability. Such a matrix is singular and has no ordinary matrix inverse. (statlect.com)

The trace is the sum of the component variances and also the sum of the eigenvalues. In principal-component analysis, this quantity represents total variance in the chosen coordinate system; its numerical value depends on variable scaling. (itl.nist.gov)

Transformations and correlation

For a deterministic matrix (A) and constant vector (b), the affine transformation (Y=AX+b) satisfies

[ \operatorname{Cov}(Y)=A\Sigma A^{\mathsf T}. ]

Adding constants therefore leaves covariance unchanged, while rescaling or mixing variables changes it predictably. This identity provides an exact rule for propagating covariance through linear transformations. (statlect.com)

If all component variances are positive, define

[ D=\operatorname{diag}(\sigma_1,\ldots,\sigma_p), \qquad \sigma_i=\sqrt{\Sigma_{ii}}. ]

The corresponding correlation matrix is

[ R=D^{-1}\Sigma D^{-1}. ]

Its diagonal entries equal one, and its off-diagonal entries are dimensionless Pearson correlations. This is the covariance matrix of variables standardized by subtracting their means and dividing by their standard deviations. The distinction matters in feature scaling, because covariance-based analyses retain the original measurement scales. (online.stat.psu.edu)

Estimation from observations

Given (n>1) independent, identically distributed observations (x_1,\ldots,x_n), let (\bar x) denote their sample mean. The conventional unbiased estimator is

[ S=\frac{1}{n-1} \sum_{k=1}^{n}(x_k-\bar x)(x_k-\bar x)^{\mathsf T}. ]

The denominator implements Bessel’s correction for estimating the mean from the same observations. If (Z) is the (n\times p) data matrix with each column centered, the equivalent expression is (S=Z^{\mathsf T}Z/(n-1)). (itl.nist.gov)

Under a multivariate Gaussian model with unknown mean, maximum likelihood estimation instead uses denominator (n). Empirical covariance estimates can be unstable when the number of variables is large relative to the number of observations, particularly when an inverse is required. (scikit-learn.org)

One form of regularization is shrinkage:

[ \widehat\Sigma_\alpha =(1-\alpha)S+\alpha\tau I, \qquad \tau=\frac{\operatorname{tr}(S)}p, \quad 0\leq\alpha\leq1, ]

where (I) is the identity matrix. This blends the empirical estimate with a spherical target, trading some estimation bias for improved stability. Robust covariance estimators address a different difficulty: sensitivity to outlying observations. (scikit-learn.org)

Statistical uses and limitations

In principal component analysis, covariance eigenvectors identify mutually orthogonal directions, and their eigenvalues give the variances along those directions. Retaining directions with the largest eigenvalues provides dimensionality reduction while preserving as much variance as possible for the retained dimension. Using a correlation matrix instead amounts to performing the analysis on standardized variables. (itl.nist.gov)

For a multivariate normal distribution, the mean vector and covariance matrix completely specify the distribution. With nonsingular covariance, its density has ellipsoidal contours determined by the quadratic form

[ (x-\mu)^{\mathsf T}\Sigma^{-1}(x-\mu). ]

This expression also accounts for different scales and correlations between coordinates. (itl.nist.gov)

Outside special distributional families, covariance does not determine the full joint distribution. Zero covariance means absence of linear association, not necessarily statistical independence. For jointly Gaussian variables, however, zero cross-covariances do imply independence. Consequently, a diagonal covariance matrix permits an independence interpretation under joint Gaussianity, but not for arbitrary data. (data140.org)