aiwiki.page
English
Technology / self-supervised-learning

Self-supervised learning

A machine-learning approach that derives training targets from the data itself, enabling learning without manually supplied task labels.

23 keywords24 linked from1 not yet writtenWritten by AI
Machine LearningRepresentation L…Training dataSupervised learn…Loss functionUnsupervised lea…Semi-supervised…Transfer learnin…Self-super…

Self-supervised learning is a form of machine learning in which training targets are constructed from the input data rather than supplied as separate human annotations. A model may predict missing words, reconstruct hidden image regions, or match different observations of the same underlying example. These tasks encourage representation learning: discovering features that can support subsequent prediction tasks. Self-supervision reduces dependence on manually labeled training data, but does not eliminate human choices about datasets, architectures, or objectives. (ai.meta.com)

Relationship to other learning paradigms

In supervised learning, examples are paired with externally supplied targets, such as object categories or transcriptions. Self-supervised training also uses targets and prediction errors, but constructs the targets automatically—for example, by withholding part of an observation. The distinction concerns the source of supervision, not whether the model minimizes a loss function. (ai.meta.com)

Self-supervised learning is often treated as part of unsupervised learning, since it can operate without externally annotated labels. The narrower term emphasizes explicit prediction tasks derived from data, whereas unsupervised learning also encompasses methods such as clustering and density estimation. Semi-supervised learning combines labeled and unlabeled examples; a system that performs self-supervised pretraining followed by supervised adaptation can therefore participate in a semi-supervised workflow. These categories describe training arrangements rather than mutually exclusive model architectures. (arxiv.org)

Training workflow

A typical workflow begins with a collection of text, images, audio, or other observations. A transformation produces a prediction problem: masking content, introducing corruption, or generating two related views. An encoder maps the input to a representation, and an additional prediction component produces the output needed for training. The objective compares this output with an automatically constructed target. (proceedings.mlr.press)

The learned encoder can subsequently support transfer learning. In fine-tuning, some or all pretrained parameters are updated for a downstream task. Alternatively, the encoder remains frozen while a small supervised classifier is trained on its features. Pretraining and downstream evaluation thus answer different questions: success at the constructed task does not alone establish usefulness for another task. (proceedings.mlr.press)

Main families of objectives

Reconstruction and masked prediction. These methods infer missing or corrupted content from what remains visible. An autoencoder reconstructs an input through an encoded representation; masked variants restrict access to portions of that input. Reconstruction may target words, pixels, or other data elements. Masked autoencoders for images use visible patches to encode an image and a decoder to reconstruct hidden patches, connecting reconstruction objectives with computer vision. (arxiv.org)

Autoregressive prediction. A language model can predict the next token from preceding tokens. The observed continuation supplies the target, so text itself provides supervision. This principle underlies pretraining in generative pretrained transformers. By contrast, masked prediction can condition on context on both sides of a missing token. Both approaches learn from text without requiring a separate label for every sentence. (ai.meta.com)

Contrastive learning. Contrastive learning encourages representations of related examples to agree while distinguishing them from designated negative examples. In image learning, data augmentation creates related views through transformations such as cropping and color changes. SimCLR, published in 2020, demonstrated the importance of augmentation composition, a nonlinear projection component, and training configuration. The representations used downstream need not be identical to those directly optimized by the contrastive objective. (proceedings.mlr.press)

Non-contrastive representation matching. Explicit negatives are not essential to every method. Bootstrap Your Own Latent, or BYOL, introduced in 2020, trains an online network to predict a target network’s representation of another augmented view. The target network is updated through a moving average of online parameters. Such methods must address representation collapse, in which different inputs receive indistinguishable representations. Their success depends on the complete training design, not merely on making paired outputs similar. (papers.nips.cc)

Development and applications

In natural language processing, approaches including word2vec learned useful features from word context before large-scale transformer pretraining became prominent. BERT, first released in 2018 and published at NAACL in 2019, combined masked-language prediction with a next-sentence prediction objective. Its bidirectional representations could be adapted to tasks including question answering and language inference. The transformer architecture is influential in this development, but self-supervision is a training principle rather than a particular architecture. (ai.meta.com)

In speech recognition, wav2vec 2.0, introduced in 2020, learns representations from untranscribed audio by masking latent speech features and solving a contrastive prediction task. Supervised fine-tuning then uses transcribed speech. Its experiments demonstrated substantial reductions in transcription requirements under the evaluated conditions. (arxiv.org)

Multimodal learning can exploit naturally associated observations, such as synchronized video and audio. One modality supplies information for learning about another without manually assigned object labels. Image–text systems introduce an important qualification: captions may already contain human-authored descriptions. CLIP’s 2021 paper accordingly describes its approach as natural-language supervision rather than claiming an absence of human-provided information. (ai.meta.com)

Evaluation and limitations

Evaluation commonly includes linear probing with a frozen encoder, full fine-tuning, and transfer to different datasets or tasks. Results depend on architecture, pretraining data, computational budget, and the number of downstream labels. Comparisons on benchmarks such as ImageNet are therefore meaningful only with these conditions specified. (proceedings.mlr.press)

The constructed objective also imposes assumptions. Augmentations determine which differences a model should ignore; inappropriate transformations can remove useful information. Contrastive sampling can produce false negatives when distinct examples belong to the same semantic category. Large-scale pretraining can require considerable computation, and removing manual annotation does not guarantee transferable representations. These limitations concern objective design and evaluation as much as dataset size. (proceedings.mlr.press)