Unsupervised learning is a branch of machine learning concerned with learning structure from observations without externally supplied target labels. Rather than learning a mapping from inputs to known answers, an algorithm may group similar observations, construct compact representations, or estimate how data are distributed. It is used in exploratory analysis and as a component of larger artificial intelligence systems. “Unsupervised” describes the source of the learning signal, not an absence of human choices about data, objectives, or model design. (deeplearningbook.org)
Learning setting and objectives
A typical dataset contains observations (x_1,\ldots,x_n), each represented by measured features. In supervised learning, these observations are paired with target values, such as categories or numerical outcomes. Unsupervised methods instead work with the observations themselves. Semi-supervised learning combines labeled and unlabeled examples. These categories describe learning settings rather than mutually exclusive families of algorithms: related mathematical techniques can operate under different forms of supervision. (deeplearningbook.org)
The distinction from self-supervised learning depends partly on terminology. Self-supervised methods construct prediction targets from the data, for example by hiding words and predicting them from their context. They share the absence of externally annotated targets with unsupervised learning, but make the internally generated supervisory task explicit. Consequently, learning from unlabeled data may be described broadly as unsupervised, while particular predictive training procedures are described more specifically as self-supervised. (ai.meta.com)
There is no single universal unsupervised objective. A loss function may measure reconstruction error, within-group dispersion, or the negative likelihood of observations. The selected objective determines which statistical relationships the model preserves. Success at optimizing that objective does not automatically establish usefulness for another task. (deeplearningbook.org)
Principal tasks and methods
Clustering. Cluster analysis organizes observations into groups according to a similarity measure or statistical model. K-means clustering partitions observations into a specified number of groups by minimizing squared distances to cluster centers. Hierarchical methods construct nested groups, while density-based methods identify dense regions and may leave some observations unassigned. These approaches embody different assumptions about cluster shape, separation, and scale; their outputs need not coincide. (sklearn.org)
Dimensionality reduction. Dimensionality reduction replaces a high-dimensional description with fewer coordinates. Principal component analysis finds orthogonal directions that successively capture the greatest remaining variance, producing a linear representation. Other methods use nonlinear relationships or different constraints. Such transformations support visualization, compression, and preprocessing, although preserving variance is not equivalent to preserving information relevant to every subsequent task. (sklearn.org)
Distribution modeling. Density estimation seeks a model of the probability distribution underlying observations. A Gaussian mixture model represents a distribution as a weighted combination of Gaussian components. It can also provide probabilistic cluster memberships rather than assigning each observation exclusively to one group. Parameters are commonly fitted through maximum likelihood estimation, often using the expectation–maximization algorithm, which alternates between estimating component memberships and updating model parameters. (scikit-learn.org)
Outlier detection. Anomaly detection identifies observations that depart from an estimated pattern of ordinary data. Unsupervised outlier detection generally assumes that unusual observations form a minority within a potentially contaminated dataset. This differs from novelty detection, in which training observations are intended to represent normal behavior and the fitted model assesses new cases. Methods may rely on local density, estimated distributions, or isolation-based partitioning. (scikit-learn.org)
Representation learning
Representation learning aims to discover useful features rather than relying entirely on manually designed ones. An autoencoder contains an encoder that transforms an input into an internal representation and a decoder that reconstructs the input. Training balances reconstruction accuracy against restrictions that prevent merely copying every observation. A narrow internal representation, sparsity constraints, or robustness to corrupted inputs can encourage the model to capture recurring structure. (deeplearningbook.org)
Autoencoders may use nonlinear artificial neural networks and can learn representations more flexible than linear projections. However, sufficient model capacity can permit memorization or an uninformative identity mapping. Regularization therefore helps define what constitutes an acceptable representation. The learned features can subsequently support classification or other tasks, but their value must be assessed beyond reconstruction performance alone. (deeplearningbook.org)
Evaluation and interpretation
Evaluation depends on the task because there may be no authoritative target answer. Clustering can be assessed through internal measures of compactness and separation. The silhouette coefficient, for example, compares an observation’s average distance within its own cluster with its average distance to the nearest alternative cluster. External evaluation is possible when reference categories are available, even if those categories were not used during training. Different criteria can favor different partitions. (scikit-learn.org)
Distribution models can be evaluated using held-out likelihood, and reconstruction models using error on unseen observations. Unsupervised methods remain susceptible to overfitting: accurately describing training observations does not guarantee generalization. When learned transformations feed into predictive evaluation, fitting those transformations on test data can leak information and distort the reported results. (deeplearningbook.org)
Data preparation and limitations
Feature engineering and preprocessing materially influence results. Distance-based methods are sensitive to measurement scales: a variable with a large numerical range can dominate a similarity calculation. Standardization changes those relative contributions, while the handling of missing values and categorical features changes the representation on which the algorithm operates. There is no preprocessing choice independent of the modeling assumptions. (scikit-learn.org)
Unsupervised outputs therefore require interpretation in relation to the data and objective. A clustering procedure supplies a partition under particular assumptions, not a unique classification of reality. Mixture models may require covariance regularization to avoid degenerate solutions, and compact representations may discard useful distinctions. Removing externally supplied labels reduces annotation requirements, but does not remove assumptions about what structure matters. (sklearn.org)