Self-attention is an attention mechanism in deep learning that computes representations of input elements using information from the same input. Each position assigns weights to accessible positions and combines their representations accordingly. “Self” identifies the shared source of the information; it does not mean that a position attends only to itself. Self-attention is a central component of the Transformer architecture, where it enables interactions between sequence elements without recurrent processing. (classic.d2l.ai)
Development and scope
Self-attention predates the Transformer. The March 2017 paper A Structured Self-attentive Sentence Embedding used attention over a sentence’s internal representations to construct a matrix-valued sentence embedding. Later that year, Attention Is All You Need introduced the Transformer for machine translation, using self-attention to compute input and output representations without sequence-aligned recurrent or convolutional layers. These designs illustrate that self-attention is a mechanism rather than a complete network architecture. (arxiv.org)
Its defining distinction from cross-attention is the source of queries, keys, and values. In self-attention, all three originate from the same sequence of representations. In an encoder–decoder architecture, cross-attention typically uses decoder representations as queries and encoder outputs as keys and values, connecting two different representation sequences. (classic.d2l.ai)
Mathematical formulation
In standard Transformer self-attention, let be a matrix containing one representation per input position. Three learned projections produce queries, keys, and values:
Although their source is identical, the resulting representations generally differ because the projection matrices differ. A query is compared with keys to determine weights; values supply the information combined into the output. (d2l.ai)
Scaled dot-product attention computes
where is the query and key dimension, denotes matrix transpose, and is an optional additive mask. Each score derives from an inner product between a query and a key. The softmax function operates independently across each row, yielding nonnegative weights that sum to one before any attention dropout. Each output row is therefore a weighted combination of value vectors. (classic.d2l.ai)
The factor controls the scale of scores as dimensionality increases. Under independent, zero-mean, unit-variance component assumptions, the unscaled dot product has variance . Scaling reduces dimension-driven softmax saturation and the associated small gradients. These assumptions motivate the factor; they are not requirements imposed on learned representations. (classic.d2l.ai)
Multiple heads and layer structure
Multi-head attention runs several attention computations with separate learned projections. Their outputs are concatenated and projected into the model’s representation dimension. This permits different heads to represent different relationships or subspaces, rather than forcing every interaction through a single set of weights. Heads are learned jointly and are not assigned fixed linguistic roles in advance. (d2l.ai)
A Transformer block contains more than attention. It also includes a positionwise feed-forward network, residual connections, and layer normalization. Attention exchanges information across positions, whereas the feed-forward network transforms each position separately. Thus, the attention matrix alone is not the block’s complete computation or its final representation. (d2l.ai)
Position and masking
Unmasked self-attention without position-dependent information is permutation-equivariant: reordering the input rows reorders the corresponding output rows. It does not, by itself, identify sequence order. Positional encoding supplies ordering information through mechanisms such as fixed sinusoidal vectors, learned position embeddings, or relative-position information. (classic.d2l.ai)
Masks restrict which positions are accessible. Padding masks exclude artificial elements introduced when batching unequal-length sequences. Causal masks exclude future positions, typically by assigning their scores negative infinity before softmax. A position can then use only itself and earlier positions. This supports left-to-right language modeling without exposing the model to future input tokens during prediction training. (d2l.smola.org)
Bidirectional attention instead permits access to both preceding and following positions. BERT, introduced in 2018, combines bidirectional Transformer representations with masked-language-model pretraining. Masking selected token identities for this training objective is distinct from applying a causal attention mask: surrounding positions remain accessible to the encoder. (arxiv.org)
Computational properties
Dense self-attention gives every position direct access to every permitted position within one layer. Unlike a recurrent neural network, it can compute representations across a supplied sequence using parallel computation. This avoids recurrent dependencies during sequence processing, although generating an autoregressive continuation still requires successive prediction steps. (classic.d2l.ai)
Its principal scaling limitation is pairwise interaction. For sequence length , the attention score matrix contains entries per head. With head dimension , dense attention’s score and value computations require arithmetic; straightforward implementations also store quadratic-size intermediates. Consequently, increasing sequence length can sharply increase computational and memory costs. (arxiv.org)
FlashAttention addresses memory traffic through tiled computation that avoids storing the full attention matrix in high-bandwidth memory. It computes exact attention, up to numerical effects, rather than approximating attention weights. This reduces memory use and transfers without eliminating dense attention’s quadratic arithmetic scaling. Sparse or approximate approaches instead modify the interaction pattern or calculation. (arxiv.org)
Applications and interpretation
Self-attention extends beyond natural language processing. In computer vision, the Vision Transformer represents an image as a sequence of embedded patches and processes it with a Transformer encoder. Attention then mixes information across image regions rather than across words. (arxiv.org)
Attention weights can be visualized, but their interpretation requires care. Experiments have found that substantially different attention distributions can produce similar predictions, while other work argues that attention’s explanatory value depends on the definition of explanation and on tests of the complete model. In explainable artificial intelligence, attention maps are therefore inspectable intermediate quantities, not automatically faithful accounts of why a prediction occurred. (aclanthology.org)