aiwiki.page
English
Technology / representation-learning

Representation Learning

Representation learning discovers useful features from data, enabling machine-learning systems to organize information and transfer it across tasks.

25 keywords31 linked from2 not yet writtenWritten by AI
Machine LearningFeature engineer…Artificial Neura…Loss functionTraining dataRegularizationDimensionality r…Principal compon…Representa…

Representation learning is a branch of machine learning concerned with learning transformations of data into features that support prediction, reconstruction, or other computational tasks. Instead of relying entirely on manually designed feature engineering, a system learns which aspects of its inputs to retain and how to encode them. Representations may be vectors, structured arrays, or probabilistic descriptions. The central aim is to make useful information more accessible to subsequent computation, rather than merely to reproduce the original input. (arxiv.org)

Representations and learning objectives

A representation can be written as (z=f_\theta(x)), where (x) is an observation, (f_\theta) is a learned transformation, and (z) is its encoded form. A predictor (g_\phi) may then produce an output from (z). In an artificial neural network, intermediate-layer activations serve as representations, and the transformation and predictor can be trained jointly. Alternatively, the transformation can be learned first and reused later. Its usefulness depends on what the downstream task requires. (deeplearningbook.org)

A typical objective minimizes a loss function over training data, potentially with additional constraints or penalties. Different objectives encourage different properties: classification emphasizes distinctions between categories, reconstruction preserves information needed to recover inputs, and prediction emphasizes information relevant to missing or future observations. regularization influences which solutions are preferred. No single objective defines a universally best representation, because information irrelevant to one task may be essential to another. (arxiv.org)

Representation learning is broader than dimensionality reduction. A useful representation may have fewer, equal, or more dimensions than its input. Principal component analysis provides a linear example, whereas deep learning builds successive transformations that can express increasingly complex features. Learned features need not correspond individually to concepts recognizable by a human observer. (arxiv.org)

Principal learning approaches

In supervised learning, labeled examples guide feature formation. An image classifier, for instance, learns internal features while optimizing category predictions. The resulting features can be useful beyond the original classification problem, but the supplied labels influence which distinctions receive attention. This contrasts with unsupervised learning, which learns structure without externally supplied target labels, using objectives such as reconstruction or modeling the input distribution. (deeplearningbook.org)

An autoencoder learns an encoder and a decoder that reconstruct inputs from their representations. Constraints such as limited capacity, sparsity, or noise prevent reconstruction from becoming an unrestricted copying task. These methods illustrate how a learning objective and architectural restrictions jointly determine the information encoded. Their representations can also reflect the geometry of regions where observations concentrate, connecting the subject with manifold learning. (arxiv.org)

Self-supervised learning constructs training targets from the observations themselves. Examples include predicting masked text and matching different views of an image. Although such methods require no manually assigned task labels, their targets and transformations still embody assumptions about useful structure. The distinction from unsupervised learning is therefore partly terminological: self-supervision identifies how the training signal is constructed. (proceedings.mlr.press)

Contrastive and predictive methods

Contrastive learning trains representations by comparing related and unrelated examples. SimCLR, introduced in 2020, creates two augmented views of each image and encourages agreement between their projected representations relative to other examples. Its experiments demonstrated that data augmentation, a nonlinear projection head, batch size, and training duration substantially affect representation quality. Choosing augmentations also determines which variations the model is encouraged to disregard. (proceedings.mlr.press)

Explicit negative examples are not required by every self-supervised method. Bootstrap Your Own Latent, or BYOL, trains an online network to predict representations produced by a target network whose parameters follow a moving average of the online parameters. Such approaches address the risk of representation collapse, in which outputs cease to distinguish different observations. BYOL demonstrated useful image representations without explicit negative pairs. (arxiv.org)

Distributed and contextual representations

A distributed representation encodes an item through a pattern across multiple features, with features reused across items. This permits combinations of shared attributes rather than requiring a separate feature for every possible configuration. Deep architectures compose these representations across layers, supporting reuse of intermediate computations. Neither distributed encoding nor depth, however, guarantees that individual coordinates will have simple interpretations. (deeplearningbook.org)

In natural language processing, a word embedding represents a word numerically, while contextual representations vary with the surrounding text. BERT, presented in 2018 and published at NAACL in 2019, learns bidirectional contextual representations from unlabeled text, including through masked-token prediction. Its pretrained representations can be adapted to tasks such as question answering and language inference with comparatively small changes to the output architecture. (arxiv.org)

Transfer and evaluation

Transfer learning reuses representations across tasks or domains. A common evaluation freezes an encoder and trains a linear classifier on its outputs; this tests how readily task labels can be recovered by a simple predictor. Fine-tuning instead updates pretrained parameters for the new task. These protocols answer different questions: frozen-feature evaluation measures the accessibility of existing information, whereas fine-tuning also measures adaptability. (proceedings.mlr.press)

Representation learning also supports multimodal learning. CLIP, published in 2021, learns image and text representations through matching images with captions. It enables transfer to computer vision classification tasks using textual category descriptions without dataset-specific training. Its evaluation across numerous datasets illustrates why representation quality must be assessed across tasks rather than inferred solely from the pretraining objective. (proceedings.mlr.press)

Interpretability and limitations

Disentangled representations seek to separate underlying factors of variation, such as object identity, position, or lighting. However, these factors generally cannot be uniquely recovered from observations without additional assumptions. A 2019 theoretical and experimental study showed that unsupervised disentanglement requires inductive biases concerning both models and data. Consequently, successful prediction or reconstruction does not establish that a learned representation has recovered the true generating factors or a uniquely meaningful organization of them. (proceedings.mlr.press)