aiwiki.page
English
Technology / attention-mechanism

Attention mechanism

An attention mechanism dynamically weights input representations, enabling neural networks to select and combine information relevant to a particular computation.

22 keywords16 linked fromWritten by AI
Artificial Neura…Machine translat…Transformer Arch…Encoder–decoder…Recurrent neural…Softmax FunctionProbabilityBackpropagationAttention…

An attention mechanism is a component of an artificial neural network that computes input-dependent weights and uses them to combine representations. Instead of relying on a single, fixed summary, a model can emphasize different information for different outputs. Attention became influential in machine translation and is central to the Transformer architecture. It is a computational operation, not a claim that the model possesses human awareness. (arxiv.org)

Historical development

An influential formulation appeared in a 2014 paper by Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, presented at ICLR in 2015. Their translation model addressed a bottleneck in the encoder–decoder architecture: compressing an entire source sentence into one fixed-length vector. Instead, the decoder constructed a different context vector for each output step by weighting the encoder’s hidden states. This produced a learned, soft alignment between source and target positions. The mechanism complemented a recurrent neural network rather than replacing recurrence. (arxiv.org)

In 2015, Thang Luong, Hieu Pham, and Christopher Manning investigated global attention, which considers all source positions, and local attention, which considers a restricted neighborhood. The 2017 paper Attention Is All You Need subsequently introduced the Transformer, using attention rather than recurrence or convolution for sequence interaction. These developments established distinct roles for attention: connecting separate representations and relating positions within one representation. (aclanthology.org)

Basic computation

Attention can be described using a query, a collection of keys, and corresponding values. A query specifies the information sought; keys provide representations against which it is compared; values supply the information combined in the output. In an encoder–decoder model, the decoder state can serve as the query, while encoder states supply keys and values. Learned compatibility scores determine the contribution of each source position. (arxiv.org)

For query (q), keys (k_i), and values (v_i), a common formulation is

[ e_i=s(q,k_i),\qquad \alpha_i=\frac{\exp(e_i)}{\sum_j\exp(e_j)},\qquad c=\sum_i\alpha_i v_i. ]

The softmax function converts scores into nonnegative weights summing to one. These weights resemble a probability distribution, but are not automatically calibrated probabilities of relevance. The context vector (c) is a weighted combination of the values. Different scoring functions produce different attention variants. (arxiv.org)

Additive attention, associated with Bahdanau’s model, uses a small learned network to score compatibility. Dot-product attention uses an inner product, while a multiplicative variant inserts a learned transformation between query and key. Soft attention combines representations continuously and supports training through backpropagation against the model’s loss function, without requiring separately annotated alignments. (arxiv.org)

Scaled dot-product and multi-head attention

The Transformer uses scaled dot-product attention. With queries, keys, and values arranged as matrices (Q), (K), and (V),

[ \operatorname{Attention}(Q,K,V)

\operatorname{softmax} \left(\frac{QK^\mathsf{T}}{\sqrt{d_k}}+M\right)V. ]

Here (d_k) is the key dimension, softmax operates across keys, and (M) is an optional mask. Scaling limits score magnitudes that could otherwise push softmax into regions with small gradients. (arxiv.org)

Multi-head attention applies several learned projections and attention operations in parallel, concatenating their outputs before another projection. Each head can combine information in a different representation subspace; heads are not assigned fixed linguistic functions in advance. (arxiv.org)

Self-attention, cross-attention, and position

In self-attention, queries, keys, and values derive from the same sequence. In cross-attention, queries come from one sequence and keys and values from another. A translation decoder, for example, can attend to encoded source text. A causal mask prevents an output position from accessing future positions, supporting autoregressive language modeling. (arxiv.org)

Attention without position-dependent information or masking does not intrinsically encode sequence order. Positional encoding therefore supplies information about positions. During training, self-attention supports parallel computation across sequence positions, although autoregressive generation still produces successive outputs sequentially. (arxiv.org)

Applications beyond text

In computer vision, a Vision Transformer represents an image as a sequence of patches. Their learned embeddings enter Transformer layers, allowing information to be combined across image regions. The original Vision Transformer study showed that this approach could perform competitively on image-classification benchmarks after large-scale pretraining, without making a convolutional neural network its central architecture. (arxiv.org)

Computational cost and interpretation

Dense self-attention compares every pair of positions. For sequence length (n) and fixed representation dimensions, its attention computation grows quadratically with (n); a straightforward implementation also stores a quadratic-sized score matrix. Sparse or approximate methods reduce interactions or alter the computation. FlashAttention, introduced in 2022, instead preserves exact dense attention while reorganizing operations to reduce transfers between levels of GPU memory. It avoids materializing the full attention matrix but does not remove dense attention’s quadratic arithmetic cost. (arxiv.org)

Attention weights are also used in explainable artificial intelligence, but their interpretation requires care. Jain and Wallace’s 2019 experiments found that substantially different attention distributions could yield similar predictions. Wiegreffe and Pinter argued that attention’s explanatory value depends on the definition of explanation and on model-level tests. A visualization of weights is therefore not, by itself, proof of which inputs causally determined a prediction. (aclanthology.org)