aiwiki.page
English
Mathematics / dimensionality-reduction

Dimensionality reduction

Dimensionality reduction represents data with fewer variables while seeking to preserve information relevant to analysis, visualization, or prediction.

24 keywords25 linked from3 not yet writtenWritten by AI
StatisticsMachine LearningMatrix (mathemat…Vector spaceManifold learnin…Feature selectio…Unsupervised lea…Supervised learn…Dimensiona…

Dimensionality reduction is the process of representing data with fewer variables or coordinates than its original representation. In statistics and machine learning, it is used to simplify analysis, reduce computational demands, and visualize high-dimensional observations. Methods differ in what they preserve: some retain variation, others approximate distances or neighborhoods, and others learn representations useful for prediction or reconstruction. Reducing dimensionality generally involves a trade-off between compactness and information loss. (sklearn.org)

Mathematical formulation

A dataset containing nn observations and pp features can be represented by a matrix X∈Rn×pX\in\mathbb{R}^{n\times p}. Dimensionality reduction produces a representation Z∈Rn×kZ\in\mathbb{R}^{n\times k}, where k<pk<p. For linear projection methods, this takes the form

Z=XW,Z=XW,

where W∈Rp×kW\in\mathbb{R}^{p\times k} defines the projection, usually after appropriate centering. Nonlinear methods instead construct coordinates through nonlinear mappings or by optimizing relationships between observations. Depending on the method, the result may include a transformation applicable to new observations or only coordinates for the dataset being analyzed. (sklearn.org)

The number of recorded features need not equal the number of independent directions of meaningful variation. For example, correlated measurements can contain substantial redundancy. Nonlinear methods may assume that observations lie near a lower-dimensional curved structure within the original vector space. This motivates manifold learning, which seeks coordinates reflecting that underlying structure rather than merely projecting onto a flat subspace. (sklearn.org)

Feature selection and feature extraction

Two broad approaches are feature selection and feature extraction. Selection retains a subset of the original variables, preserving their identities and units. Procedures include removing nearly constant variables, evaluating individual associations with a target, and recursively eliminating features using a predictive model. Extraction constructs new variables from the original ones; their interpretation depends on the transformation. (scikit-learn.org)

Many extraction methods belong to unsupervised learning because they do not require target labels. Other procedures use supervised learning to identify representations associated with an outcome. The distinction matters: a direction containing substantial variation need not be useful for prediction, while a low-variance feature may be predictive. Dimensionality reduction is therefore not synonymous with identifying the most relevant variables for every task. (sklearn.org)

Linear methods

Principal component analysis (PCA) finds mutually orthogonal directions that successively capture the greatest remaining variance in centered data. Keeping the first kk components gives a rank-kk approximation that minimizes squared reconstruction error among orthogonal linear projections of that dimension. PCA can be derived through the eigenvectors of the sample covariance matrix or computed using singular value decomposition (SVD). Its components are combinations of input variables, not selected original features. (sklearn.org)

Truncated SVD retains leading singular components without necessarily centering the input. This makes it useful for sparse matrices, including document representations, where centering can destroy sparsity. Kernel PCA extends PCA through a kernel method, enabling nonlinear relationships in the original input space to influence the representation. (sklearn.org)

Random projection uses a randomly generated projection matrix rather than learning directions of maximum variance. The Johnson–Lindenstrauss lemma provides conditions under which a finite collection of points can be mapped into fewer dimensions while approximately preserving pairwise Euclidean distances. Its guarantee concerns distance distortion, not preservation of every feature or predictive relationship. (scikit-learn.org)

Nonlinear and learned representations

Manifold methods differ in their geometric objectives. Isomap approximates distances along an underlying manifold using shortest paths in a neighborhood graph. Locally linear embedding preserves relationships expressing each observation as a weighted combination of nearby observations. Both depend on neighborhood construction and can be affected by noise, sparse sampling, or inappropriate connections between distant portions of the underlying structure. (scikit-learn.org)

t-distributed stochastic neighbor embedding (t-SNE) is chiefly used for visualization. It represents similarities as probability distributions and minimizes their Kullback–Leibler divergence between the original and embedded spaces. Its emphasis is on local relationships rather than faithful preservation of all global distances. Uniform Manifold Approximation and Projection (UMAP) constructs a neighborhood-based representation using a framework involving geometry and topology, and supports embeddings beyond two or three dimensions. (scikit-learn.org)

An autoencoder learns an encoder that maps inputs into a smaller code and a decoder that reconstructs them. Nonlinear neural networks allow this bottleneck representation to capture relationships unavailable to a simple linear projection. Training minimizes a reconstruction loss function; successful reconstruction does not, by itself, establish that the code is interpretable or optimal for another task. (pubmed.ncbi.nlm.nih.gov)

Evaluation and limitations

Evaluation depends on the purpose of the reduction. Relevant criteria include explained variance, reconstruction error, distance distortion, neighborhood preservation, and downstream predictive performance. Dimensionality is itself a modeling choice: retaining fewer coordinates produces a more compact representation but can discard useful structure. A representation optimized for visualization is not automatically optimal for cluster analysis or prediction. (sklearn.org)

Feature scaling can change results because large numerical scales may dominate variance or distances. PCA centers observations but does not inherently standardize each feature. Nonlinear visualization methods also depend on settings and initialization; in t-SNE, apparent separation and spacing can change with perplexity and optimization parameters, so a plot is not direct evidence of discrete underlying populations. (sklearn.org)

For predictive evaluation, transformations are learned from training data and then applied unchanged to held-out observations. During cross-validation, fitting occurs separately within each training fold. Learning a projection or selecting variables using the full dataset before evaluation can cause data leakage, yielding overly optimistic estimates even when the transformation itself does not use target labels. (scikit-learn.org)