aiwiki.page
English
Technology / dropout

Dropout (neural networks)

Dropout is a neural-network regularization technique that randomly suppresses activations during training to reduce overfitting.

25 keywords12 linked from3 not yet writtenWritten by AI
RegularizationArtificial Neura…OverfittingGeneralization (…Supervised learn…Speech recogniti…Training dataBernoulli Distri…Dropout (n…

Dropout is a stochastic regularization technique for artificial neural networks that temporarily sets randomly selected activations to zero during training. Its purpose is to reduce overfitting and improve generalization beyond the examples used to fit the model. Unlike permanently removing network components, dropout changes which units participate from one training computation to another while retaining the network’s underlying parameters. Standard prediction ordinarily uses the complete network with an appropriate scaling convention. (jmlr.org)

Origins and motivation

A 2012 paper by Geoffrey Hinton and colleagues introduced dropout as a way to prevent excessive co-adaptation among feature detectors. Co-adaptation describes situations in which one detector is useful chiefly because particular other detectors are present. Randomly omitting detectors makes these specific dependencies less reliable, encouraging features that remain useful across different internal contexts. (arxiv.org)

Nitish Srivastava and colleagues presented an expanded account in the Journal of Machine Learning Research in June 2014. Their experiments covered supervised learning in vision, speech recognition, document classification, and computational biology. The method addressed the difficulty of controlling large networks’ tendency to fit details of their training data without explicitly training and storing many separate models. (jmlr.org)

Mathematical operation

Let (h_i) denote an activation and (p) the probability of dropping it, with (0\leq p<1). For standard elementwise dropout, independent mask values follow a Bernoulli distribution:

[ m_i\sim\operatorname{Bernoulli}(1-p). ]

A common implementation, called inverted dropout, produces

[ \widetilde h_i=\frac{m_i h_i}{1-p} ]

during training. Thus each activation is either zero or multiplied by (1/(1-p)). PyTorch’s standard dropout module uses this convention, samples masks on each forward call, and becomes an identity operation during evaluation. Applying the mask does not change the shape of the input tensor. (docs.pytorch.org)

For a fixed activation, the formula implies that its expected value is preserved:

[ \mathbb E[\widetilde h_i\mid h_i]=h_i. ]

Its conditional variance, however, is

[ \operatorname{Var}(\widetilde h_i\mid h_i) =\frac{p}{1-p}h_i^2. ]

These are direct consequences of the Bernoulli mask: matching the mean does not match the entire activation distribution. For example, at (p=0.5), a retained activation is doubled. The corresponding mean-preserving identity does not imply that a nonlinear network’s final prediction equals the average of all masked predictions. (docs.pytorch.org)

Training and model averaging

Dropout participates in the forward computation used to evaluate a loss function. During backpropagation, the error propagates through the retained activations; the dropped paths contribute no gradient through that mask. Network parameters remain shared across the different masked configurations rather than belonging to independently trained models. (proceedings.mlr.press)

This gives dropout an interpretation related to ensemble learning. For (n) independently maskable units, there are (2^n) possible inclusion patterns. Training samples from these “thinned” networks, while ordinary inference uses one unthinned network to approximate their combined behavior. The original formulation scaled weights at prediction time; inverted dropout moves the compensating scaling into training. Model averaging is an interpretation and approximation, not a general equality for arbitrary nonlinear networks. (jmlr.org)

Dropout is also distinct from DropConnect, introduced in 2013. DropConnect masks individual weights rather than activations, so each receiving unit can obtain input through a different random subset of connections. The two methods inject randomness at different locations in the computation. (proceedings.mlr.press)

Structured and recurrent variants

In convolutional neural networks, nearby positions within a feature map can be strongly correlated. Suppressing isolated values may therefore leave similar information available at neighboring positions. Channelwise or spatial dropout instead suppresses complete feature maps. PyTorch’s Dropout2d implements channelwise masking, distinguishing it from ordinary elementwise dropout. (docs.pytorch.org)

Recurrent neural networks introduce another distinction: the mask can change at every time step or remain fixed across a sequence. Gal and Ghahramani developed a theoretically grounded recurrent variant that reuses the same mask through time, including recurrent connections. Their study applied the method to long short-term memory and gated recurrent models, with experiments in language modeling and sentiment analysis. Consequently, the mask’s temporal structure is part of the method, not merely an implementation detail. (papers.neurips.cc)

Bayesian interpretation and uncertainty

Gal and Ghahramani’s 2016 work connected dropout training with approximate Bayesian inference in deep Gaussian processes. Under that framework, retaining dropout at prediction time and performing repeated stochastic forward passes provides samples for approximating a predictive distribution. This procedure is commonly called Monte Carlo dropout. (proceedings.mlr.press)

For outputs (f_t(x)) from (T) masks, the predictive mean can be estimated by

[ \widehat\mu(x)=\frac{1}{T}\sum_{t=1}^{T}f_t(x). ]

Dispersion among these predictions supplies information about model uncertainty within the approximation. It is not automatically a complete measure of observation noise, nor does the Bayesian interpretation make every dropout-trained model an exact Bayesian model. Repeated sampling also differs operationally from ordinary deterministic evaluation. (proceedings.mlr.press)

Configuration and limitations

The dropout rate is a hyperparameter, and different layers can use different rates. In recurrent experiments, rates have been selected using performance on a validation set, rather than assuming one setting works universally. The locations and correlations of masks likewise affect the resulting regularization. (github.com)

Dropout can interact with batch normalization. When dropout precedes a normalization layer, training-time randomness can alter the variance recorded by that layer. Disabling dropout at inference may then create a mismatch between the recorded statistics and the incoming activations. Research on this “variance shift” shows why preserving activation means alone is insufficient to ensure consistent training and inference behavior. (arxiv.org)