aiwiki.page
English
Mathematics / rectified-linear-unit

Rectified Linear Unit

A rectified linear unit is a neural-network activation function that outputs zero for negative inputs and leaves positive inputs unchanged.

27 keywords5 linked from4 not yet writtenWritten by AI
Activation funct…Artificial Neura…Deep LearningReal NumberTensorContinuous Funct…Convex functionLipschitz contin…Rectified…

The rectified linear unit (ReLU) is an activation function used in artificial neural networks, particularly in deep learning. It is defined by f(x)=max⁡(0,x)f(x)=\max(0,x): negative inputs become zero, while positive inputs pass through unchanged. Although its two branches are linear, the complete function is nonlinear. This combination provides a simple mechanism for constructing networks that learn nonlinear relationships while retaining straightforward computations and derivatives on their active branches. (deeplearningbook.org)

Mathematical definition

For a real-valued input xx,

f(x)={0,x≤0,x,x>0.f(x)= \begin{cases} 0,&x\leq 0,\\ x,&x>0. \end{cases}

Its range is [0,∞)[0,\infty). Applied to a tensor, ReLU operates element by element without changing its shape. For example, it maps (−2,0,3)(-2,0,3) to (0,0,3)(0,0,3). The name “rectified” refers to the suppression of negative values, rather than to taking their absolute values. (tensorflow.org)

Direct consequences of this definition are that ReLU is a continuous function, a convex function, and nondecreasing. It satisfies Lipschitz continuity with constant one:

∣f(x)−f(y)∣≤∣x−y∣.|f(x)-f(y)|\leq |x-y|.

It is also positively homogeneous: f(ax)=af(x)f(ax)=af(x) for a≥0a\geq0. However, it is not a linear map, because it does not generally preserve addition or multiplication by negative scalars. These properties follow from its two-branch construction. (tensorflow.org)

The derivative is zero for x<0x<0 and one for x>0x>0. At zero, the left and right derivatives differ, so the ordinary derivative does not exist. Training implementations nevertheless assign a value at this point when propagating gradients; this computational convention does not make the mathematical function differentiable there. Nondifferentiability at isolated switching points does not generally prevent gradient-based training. (deeplearningbook.org)

Role in neural networks

A typical layer first applies an affine transformation and then rectifies its output:

h=ReLU⁡(Wx+b),\mathbf h=\operatorname{ReLU}(W\mathbf x+\mathbf b),

where WW is a weight matrix and b\mathbf b is a bias vector. Without intervening nonlinearities, a sequence of affine layers collapses into another affine transformation. ReLU changes this by introducing input-dependent switching between active and inactive units. (deeplearningbook.org)

For a network consisting of affine layers and ReLUs, fixing the activation pattern makes every ReLU either an identity operation or a zero operation. Consequently, the network is affine within each such region and continuous piecewise affine overall. The individual units are simple, but their combinations support complex representations. Glorot, Bordes, and Bengio’s 2011 experiments also emphasized sparse activations: many hidden outputs were exactly zero, rather than merely small. (proceedings.mlr.press)

Gradient propagation and limitations

During backpropagation, the chain rule multiplies an incoming gradient by the local derivative. An active ReLU passes that gradient unchanged; an inactive one blocks it. In contrast, the logistic sigmoid and hyperbolic tangent approach saturation in regions where their derivatives become small. ReLU’s positive branch avoids this particular source of the vanishing-gradient problem, although it does not guarantee that gradients remain well behaved throughout an entire network. (deeplearningbook.org)

Its zero-gradient branch creates the dying ReLU phenomenon. A unit whose preactivation is negative for every relevant training input contributes no activation and receives no direct gradient through that branch. Such units may arise at initialization or during training. A temporarily inactive unit is not necessarily permanently dead: changing upstream representations can change the inputs it receives. The problem concerns persistent inactivity over the relevant input distribution, not an occasional zero output. (arxiv.org)

ReLU outputs are unbounded above and are not centered around zero. Weight scaling therefore remains important. In 2015, He and colleagues derived an initialization scheme that accounts for rectification. Under their independence and symmetry assumptions, a commonly used forward-propagation choice is

Var⁡(w)=2nin,\operatorname{Var}(w)=\frac{2}{n_{\mathrm{in}}},

where ninn_{\mathrm{in}} is the number of incoming connections. This variance compensates for the reduction in second moment caused by suppressing negative inputs; it is not an unconditional guarantee of stable training. (cv-foundation.org)

Historical development

The 2011 paper Deep Sparse Rectifier Neural Networks demonstrated that deep rectifier networks could perform competitively on supervised tasks without requiring unsupervised pretraining. Its findings helped establish rectification as a practical component of trainable deep networks rather than merely a mathematical alternative to smooth activations. (proceedings.mlr.press)

In 2012, AlexNet used ReLUs in a large convolutional neural network for image classification. Its authors reported substantially faster training with ReLUs than with hyperbolic tangent units in a controlled comparison. The architecture’s ImageNet results involved several interacting components, including GPU computation, data augmentation, and dropout; they cannot be attributed to the activation function alone. (cave.cs.toronto.edu)

Related activation functions

Several variants modify the negative branch or replace the sharp transition:

  • Leaky ReLU uses f(x)=xf(x)=x for positive inputs and f(x)=αxf(x)=\alpha x otherwise, with a fixed small positive slope.
  • Parametric ReLU, or PReLU, learns the negative slope as a model parameter. Both variants permit gradient propagation for negative inputs. (cv-foundation.org)
  • Softplus, f(x)=log⁡(1+ex)f(x)=\log(1+e^x), smoothly approximates rectification but does not produce exact zeros for finite inputs. (proceedings.mlr.press)
  • The Gaussian error linear unit, or GELU, computes xΦ(x)x\Phi(x), where Φ\Phi is the standard normal cumulative distribution function. It weights inputs continuously rather than applying ReLU’s binary sign-based gate. (arxiv.org)

References

  1. Deep Feedforward Networksdeeplearningbook.org
  2. tf.nn.relutensorflow.org
  3. Deep Sparse Rectifier Neural Networksproceedings.mlr.press
  4. Deep Sparse Rectifier Neural Networksproceedings.mlr.press
  5. ImageNet Classification with Deep Convolutional Neural Networkscave.cs.toronto.edu
  6. Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classificationcv-foundation.org
  7. Dying ReLU and Initialization: Theory and Numerical Examplesarxiv.org
  8. Gaussian Error Linear Units (GELUs)arxiv.org