The multivariate normal distribution is a joint probability distribution that generalizes the normal distribution to several variables considered simultaneously. Its defining property is that every linear combination of its components has a univariate normal distribution, allowing constant variables as zero-variance cases. A multivariate normal law is completely determined by a mean vector and a covariance matrix. It provides a mathematical framework for describing correlated measurements and for deriving multivariate statistical procedures. (www2.stat.duke.edu)
Definition and parameters
Let (X=(X_1,\ldots,X_d)^\mathsf{T}) be a vector of real-valued random variables. The notation [ X\sim\mathcal N_d(\mu,\Sigma) ] means that, for every (a\in\mathbb R^d), [ a^\mathsf{T}X\sim \mathcal N(a^\mathsf{T}\mu,;a^\mathsf{T}\Sigma a). ] Here (\mu=\mathbb E[X]) is the vector of component expected values, and [ \Sigma=\mathbb E[(X-\mu)(X-\mu)^\mathsf{T}]. ] The superscript (\mathsf{T}) denotes matrix transpose. The diagonal entries of (\Sigma) are component variances, while its off-diagonal entries are pairwise covariances. The matrix is symmetric and positive semidefinite. (www2.stat.duke.edu)
An equivalent construction is (X=\mu+AZ), where (Z) contains independent standard normal variables and (AA^\mathsf{T}=\Sigma). This construction includes singular covariance matrices. Importantly, normality of each component separately does not establish joint normality: the definition concerns all linear combinations, not merely individual coordinates. (live.ocw.mit.edu)
Density and geometric interpretation
When (\Sigma) is positive definite, the probability density function is [ f_X(x)= \frac{1}{(2\pi)^{d/2}(\det\Sigma)^{1/2}} \exp!\left[-\frac12 (x-\mu)^\mathsf{T}\Sigma^{-1}(x-\mu)\right]. ] The determinant controls the normalization, while the inverse matrix determines how departures from the mean are weighted. The density has its unique maximum at (\mu), and its constant-density surfaces are ellipsoids centered there. In two dimensions, these contours are ellipses. (ocw.mit.edu)
The quadratic expression [ D^2(x)=(x-\mu)^\mathsf{T}\Sigma^{-1}(x-\mu) ] is the squared Mahalanobis distance. Unlike ordinary distance, it accounts for both unequal scales and dependence between coordinates. The eigenvectors of (\Sigma) specify the ellipsoids’ principal directions; their semiaxis lengths are proportional to the square roots of the corresponding eigenvalues. (alumni.media.mit.edu)
If (\Sigma) is singular, the distribution remains well defined but has no density with respect to (d)-dimensional volume. It is concentrated on an affine subspace whose dimension equals the rank of (\Sigma). Such degeneracy represents exact linear constraints among components, rather than an invalid probability model. (ocw.mit.edu)
Transformations and independence
Multivariate normality is preserved by an affine transformation. For a fixed matrix (B) and vector (b), [ BX+b\sim \mathcal N(B\mu+b,;B\Sigma B^\mathsf{T}). ] Consequently, selecting any subset of coordinates gives a multivariate normal marginal distribution, with the corresponding entries of (\mu) and principal submatrix of (\Sigma). Linear projections therefore remain Gaussian even when they reduce dimension. (www2.stat.duke.edu)
For jointly Gaussian variables, zero covariance implies statistical independence. In particular, a diagonal covariance matrix means that all components are mutually independent. More generally, two Gaussian subvectors are independent exactly when their cross-covariance matrix is zero. This implication is special: uncorrelated variables need not be independent outside the jointly Gaussian setting. (live.ocw.mit.edu)
Conditional distributions
Partition the vector and its parameters as [ X=\begin{pmatrix}X_1\X_2\end{pmatrix},\qquad \mu=\begin{pmatrix}\mu_1\\mu_2\end{pmatrix},\qquad \Sigma= \begin{pmatrix} \Sigma_{11}&\Sigma_{12}\ \Sigma_{21}&\Sigma_{22} \end{pmatrix}. ] If (\Sigma_{22}) is invertible, the conditional distribution of (X_1) given (X_2=x_2) is normal, with parameters [ \mu_{1\mid2} =\mu_1+\Sigma_{12}\Sigma_{22}^{-1}(x_2-\mu_2), ] [ \Sigma_{1\mid2} =\Sigma_{11}-\Sigma_{12}\Sigma_{22}^{-1}\Sigma_{21}. ] The conditional expectation is thus affine in the observed value, whereas the conditional covariance does not depend on that value. These formulas explain why Gaussian conditioning produces linear predictions together with an explicit remaining uncertainty. (online.stat.psu.edu)
For example, if two components have means zero, variances one, and correlation (\rho), with (|\rho|<1), then [ X_1\mid X_2=x\sim\mathcal N(\rho x,;1-\rho^2). ] This is the scalar specialization of the block formulas: observing one component shifts the predicted mean of the other and reduces its variance. (online.stat.psu.edu)
Simulation and statistical uses
Simulation follows directly from (X=\mu+AZ). For positive definite (\Sigma), a Cholesky decomposition supplies a triangular factor (L) satisfying (\Sigma=LL^\mathsf{T}); drawing independent standard normals and computing (\mu+LZ) generates the desired sample. An eigenvalue decomposition provides a construction for positive semidefinite, including singular, covariance matrices. (live.ocw.mit.edu)
The multivariate central limit theorem supplies another reason for the distribution’s importance. For independent, identically distributed random vectors with mean (\mu) and finite covariance (\Sigma), [ \sqrt n(\bar X_n-\mu) \xrightarrow{\mathrm d}\mathcal N_d(0,\Sigma). ] Thus appropriately scaled sample means approach a Gaussian law even when the original observations are not Gaussian. (ocw.mit.edu)
In machine learning, a Gaussian process is defined by requiring every finite collection of its values to have a multivariate normal distribution. The conditional formulas then yield predictive distributions for unobserved values from observed data and a specified covariance function. (see.stanford.edu)