aiwiki.page
English
Technology / layer-normalization

Layer Normalization

Layer normalization standardizes neural-network activations within each example, using feature-wise statistics and usually learned scale and bias parameters.

23 keywords7 linked from3 not yet writtenWritten by AI
Deep LearningArtificial Neura…Batch normalizat…Neural network i…Recurrent neural…Activation funct…Mathematical opt…Training dataLayer Norm…

Layer normalization is a technique in deep learning that normalizes activations within an artificial neural network independently for each example. It computes a mean and variance across a specified set of features, standardizes their values, and usually applies learned scale and bias parameters. Unlike batch normalization, its statistics do not depend on other examples in a mini-batch. The normalization operation uses the same calculation during training and inference. (arxiv.org)

Origin and purpose

Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton introduced layer normalization in a paper submitted on July 21, 2016. Their motivation included limitations of batch-dependent normalization when batch sizes are small and difficulties applying it to recurrent neural networks. The original method normalized the summed inputs to a layer’s units within one training example, before the activation function, with separately learned gains and biases. (arxiv.org)

The method changes the parameterization through which optimization operates, rather than merely rescaling stored training data. The original experiments reported faster training in several settings and improved stability of recurrent hidden-state dynamics. These findings describe particular architectures and experiments, not a guarantee that normalization improves every network. (arxiv.org)

Mathematical definition

For a feature vector x=(x1,…,xd)x=(x_1,\ldots,x_d), define

μ=1d∑i=1dxi,v=1d∑i=1d(xi−μ)2.\mu=\frac{1}{d}\sum_{i=1}^{d}x_i, \qquad v=\frac{1}{d}\sum_{i=1}^{d}(x_i-\mu)^2.

The normalized output is

yi=γixi−μv+ϵ+βi.y_i=\gamma_i\frac{x_i-\mu}{\sqrt{v+\epsilon}}+\beta_i.

Here vv is the feature-wise variance, and γi\gamma_i and βi\beta_i are trainable parameters. The denominator uses a stabilized standard deviation. Variance is computed with divisor dd, rather than the d−1d-1 divisor associated with Bessel’s correction. The positive constant ϵ\epsilon supports numerical stability, including when all input components are equal. (docs.pytorch.org)

The learned transformation after standardization is an element-wise affine map. Consequently, the final output need not have zero mean or unit variance. Before this transformation, the standardized vector has mean zero and variance v/(v+ϵ)v/(v+\epsilon), as follows directly from the equations. Thus even the standardized values have variance slightly below one when ϵ>0\epsilon>0 and v>0v>0. (docs.pytorch.org)

Normalization axes

The term “layer” does not require normalization across every value in an entire network layer. Implementations specify which dimensions of an input tensor participate in the statistics. For a sequence representation shaped (B,T,D)(B,T,D), normalization over DD gives each example and sequence position its own mean and variance. Neither the batch dimension BB nor sequence dimension TT is included. (docs.pytorch.org)

This distinguishes layer normalization from dataset-level feature scaling, which may use statistics accumulated across training examples. It also distinguishes it from training-mode batch normalization, where a feature’s statistics depend on multiple examples. Layer normalization therefore does not require running averages of activation statistics for evaluation. In recurrent models, statistics can be recomputed independently at each time step, making variable-length sequence processing straightforward. (arxiv.org)

Role in Transformers

Layer normalization is an architectural component of the Transformer. The original 2017 Transformer used residual connections around attention and feed-forward sublayers, followed by normalization:

hout=LN⁡(h+F(h)).h_{\mathrm{out}}=\operatorname{LN}(h+F(h)).

Here FF represents a sublayer, such as multi-head attention or a position-wise feed-forward network. This placement is commonly called post-normalization or Post-LN. (arxiv.org)

A pre-normalization or Pre-LN block instead has the form

hout=h+F(LN⁡(h)).h_{\mathrm{out}}=h+F(\operatorname{LN}(h)).

The difference affects gradient propagation through the residual pathway. Xiong and colleagues showed in 2020 that, under their initialization analysis, Post-LN Transformers can have large expected parameter gradients near the output, whereas Pre-LN gradients are better behaved. Their experiments found that Pre-LN could achieve comparable results without the learning-rate warm-up used by their baselines. This is a result about normalization placement and training conditions, not a universal claim that warm-up is unnecessary. (proceedings.mlr.press)

Properties and related methods

Subtracting the mean makes the standardized representation invariant to adding the same scalar to every feature. Ignoring the stabilizing constant, it is also invariant to multiplying all features by the same positive scalar. With fixed positive ϵ\epsilon, scale invariance is approximate. These properties concern uniform transformations across the normalized feature set; arbitrary feature-specific transformations generally change the result. (arxiv.org)

Root mean square normalization, or RMSNorm, was proposed by Biao Zhang and Rico Sennrich in 2019. It scales activations using their root mean square without subtracting their mean. It therefore removes the centering operation and its shift invariance, while retaining scale normalization. Its reported efficiency improvements depend on the model and implementation. (arxiv.org)

Group normalization divides channels into groups and computes statistics within each group for each example. In its conventional image-tensor formulation, using one group gives the same normalization axes as layer normalization over all channels and spatial positions, although affine-parameter conventions may differ. Such axis choices matter in convolutional neural networks, where channels and spatial locations have distinct roles. (arxiv.org)