Batch normalization is a technique used in deep learning to normalize intermediate activations within an artificial neural network. During training, it uses statistics computed from a mini-batch, then applies learned scaling and shifting parameters. Introduced by Sergey Ioffe and Christian Szegedy in 2015, it can accelerate training and permit higher learning rates. Unlike ordinary input preprocessing, normalization is integrated into the model itself. (proceedings.mlr.press)
Mathematical operation
For one feature, let denote the values included in a normalization batch. The batch mean and variance are
The layer computes
Here, is a numerical-stability constant, while and are trainable scale and offset parameters. They allow the network to select an appropriate output scale and mean rather than permanently enforcing standardized outputs. Common initialization sets and . (docs.pytorch.org)
The operation is differentiable, and its dependence on batch statistics is included in backpropagation. Consequently, training examples are coupled: changing one example can change the normalized activations of other examples in the same batch. Batch normalization does not perform full whitening or force activations to follow a normal distribution; it standardizes individual features rather than eliminating all correlations between them. (proceedings.mlr.press)
This differs from conventional feature scaling, where fixed statistics are estimated from training data before model fitting. A preprocessing normalization layer applies those stored statistics to inputs, whereas batch normalization recomputes statistics during training and includes learned affine parameters. (keras.io)
Placement and normalization axes
In a fully connected network, statistics are generally computed separately for each feature across the mini-batch. The original formulation placed normalization between a linear transformation and its activation function. This makes normalization part of the network’s computation rather than a separate transformation of the dataset. (proceedings.mlr.press)
For a convolutional neural network, normalization usually operates separately on each channel. Given a tensor with shape , the mean and variance for a channel are computed over the batch and spatial dimensions . Each channel has its own scale and offset, shared across spatial positions. Thus, the number of values contributing to a channel’s statistics can exceed the number of images in the batch. (docs.pytorch.org)
Training and inference
Batch normalization normally has distinct training and inference behavior. Training uses current-batch statistics while also updating moving estimates of the mean and variance. Inference substitutes those stored estimates:
With fixed running statistics, an example’s output no longer depends on the other examples processed alongside it. The running statistics are state variables, not parameters optimized through gradients. Their suitability depends on how closely the inference data resemble the data used to estimate them. (keras.io)
Implementation details differ. PyTorch uses the variance estimator with denominator for the training forward pass, but an unbiased estimator for updating running variance. It also permits batch statistics at evaluation when running-statistic tracking is disabled. (docs.pytorch.org)
The parameter named “momentum” controls the moving average, not optimizer momentum. PyTorch weights the new statistic by this parameter; Keras weights the previous running statistic by it. Identical numerical settings therefore do not imply identical updates across these frameworks. (docs.pytorch.org)
Effects on optimization and generalization
The original paper motivated batch normalization through internal covariate shift: changes in intermediate activation distributions as earlier layers’ parameters are updated. Its experiments showed faster convergence, tolerance of larger learning rates, and reduced sensitivity to initialization. Batch-dependent fluctuations also provided a regularization effect, sometimes reducing the need for dropout. These findings were experimental results, not guarantees for every architecture or task. (proceedings.mlr.press)
Later research questioned whether reducing distributional change adequately explains the technique’s effectiveness. A 2018 study by Santurkar and colleagues instead emphasized smoother optimization behavior: under their analyses and experiments, batch normalization made the loss and gradients more predictable under parameter changes. This offers an explanation for more stable, faster optimization without treating internal covariate shift as the sole cause. (papers.nips.cc)
Optimization and generalization are separate effects. Other empirical research found that normalization prevented uncontrolled activation growth under large gradient updates, enabling larger steps and potentially helping training avoid sharp minima. Its influence therefore extends beyond merely rescaling the forward-pass values. (arxiv.org)
Limitations and related methods
Small batches can produce inaccurate statistics and impair performance. This is particularly relevant in computer vision tasks such as semantic segmentation, where large inputs constrain batch size through memory requirements. The group-normalization paper demonstrated substantial batch-size sensitivity for batch-normalized models in its experiments. (ecva.net)
In parallel training, synchronized batch normalization computes statistics across participating devices rather than independently within each device’s local batch. It changes the set of activations contributing to normalization without changing the basic scale-and-shift operation. (docs.pytorch.org)
Layer normalization computes statistics within an individual example across its features, avoiding dependence on other batch members. Its original formulation uses the same computation during training and testing and was developed partly to address difficulties applying batch normalization to recurrent neural networks. (arxiv.org)
Group normalization divides channels into groups and normalizes within each example over each group’s channels and spatial positions. Its statistics are independent of batch size, distinguishing it from both batch normalization and feature-wide layer normalization. (ecva.net)