Principal component analysis (PCA) is a method in statistics and linear algebra that represents multivariate data through mutually orthogonal directions of greatest variation. It transforms the original variables into linear combinations called principal components, ordered by the variance they explain. Retaining only the leading components provides dimensionality reduction. In machine learning, ordinary PCA is an unsupervised learning method because its directions are determined without target labels. (arxiv.org)
Geometric interpretation and origins
PCA treats observations as points in a vector space. After subtracting the mean of each variable, the first principal direction is the axis along which projected observations have the largest variance. The second maximizes the remaining projected variance subject to being perpendicular to the first; subsequent directions follow the same rule. The resulting component scores are uncorrelated, although this does not generally make them statistically independent. (arxiv.org)
An equivalent interpretation is to find a lower-dimensional linear subspace minimizing the sum of squared perpendicular distances from the centered observations. Karl Pearson described this closest-fitting-lines-and-planes formulation in 1901. Unlike ordinary linear regression, which distinguishes a response from predictors, this geometric formulation treats the measured coordinates symmetrically, for the chosen scaling. (pca.narod.ru)
For illustration, an elongated cloud of two-dimensional points has a first principal direction along its long axis. Projection onto that axis preserves more variation than projection onto any other single direction; the perpendicular second direction describes its width. (arxiv.org)
Mathematical formulation
Let be an centered data matrix, with observations in rows and variables in columns, where . Its sample covariance matrix is
Because is symmetric and positive semidefinite, it has orthonormal eigenvectors and nonnegative eigenvalues ordered as . The first direction solves
The solution is an eigenvector associated with . Later directions solve the corresponding constrained problems with orthogonality to earlier directions. (arxiv.org)
For , the reduced coordinates, or scores, are . Their sample covariance is diagonal, with entries . The reconstruction in centered coordinates is
Restoring the original column means returns the approximation to the original coordinate system. This projection minimizes squared reconstruction error among -dimensional orthogonal projections. Thus variance maximization and a reconstruction-error loss function describe the same solution. (arxiv.org)
Computation and interpretation
PCA can be computed directly using the singular value decomposition:
The right singular vectors give the principal directions, while gives their scores. If is the corresponding singular value, then . Centered data have at most nonzero components. Singular value decomposition therefore connects PCA with low-rank matrix approximation and provides a computational route without explicitly constructing the covariance matrix. (arxiv.org)
The coefficients defining a component are often called loadings, although some conventions reserve that term for eigenvectors multiplied by square roots of eigenvalues. Scores describe observations; loadings describe how variables contribute. Reversing a direction’s sign reverses its scores but leaves the representation unchanged. Repeated eigenvalues permit different orthonormal bases for the same eigenspace, so individual directions within it are not uniquely determined. (arxiv.org)
Scaling and component selection
Centering and scaling are distinct operations. Covariance-based PCA retains the variables’ original scales, so changing measurement units can change its directions. Standardizing nonconstant variables to unit variance instead yields correlation-based PCA. This prevents large numerical scales from automatically dominating, but also changes the question from absolute variation to variation relative to each variable’s spread. (arxiv.org)
Provided total variance is positive, component explains the fraction
Cumulative explained variance measures how much variation the first components retain. A scree plot displays the eigenvalues against component number, helping identify where their decline flattens. Neither a visual elbow nor a fixed variance threshold universally determines the appropriate dimension; selection depends on the analytical objective. (arxiv.org)
For predictive evaluation, means, scales, and directions are estimated from training data alone. The same fitted transformation is then applied to held-out observations. During cross-validation, preprocessing is fitted separately inside each training fold; otherwise information from evaluation observations can cause data leakage. (scikit-learn.org)
Applications, limitations, and variants
PCA supports visualization, compression, noise reduction, and exploratory analysis of high-dimensional measurements, including gene-expression data. Unlike feature selection, it constructs combinations of variables rather than selecting a subset of original measurements. Its usefulness depends on whether dominant variation corresponds to the structure of interest. Large variance need not imply predictive relevance, and ordinary PCA does not optimize separation between target classes. (arxiv.org)
PCA captures linear structure and can be influenced strongly by outliers. Its components describe statistical variation, not necessarily causal mechanisms. Related methods alter these assumptions: kernel PCA uses a kernel method to obtain nonlinear representations; sparse PCA encourages many coefficients to be zero; and probabilistic PCA introduces a latent-variable model with isotropic Gaussian noise. Incremental algorithms process batches, while randomized decompositions approximate leading components when a full decomposition would be expensive. (arxiv.org)