An activation function is a mathematical function that determines the output of a unit or layer in an artificial neural network. It usually acts on a weighted combination of inputs plus a bias. Nonlinear activations enable networks to represent relationships that cannot be expressed by affine transformations alone; output activations can constrain predictions to particular ranges or probability distributions. Their mathematical properties influence both representational capacity and the behavior of learning in deep learning. (deeplearningbook.org)
Mathematical formulation and role
For a typical unit, the computation is
where are inputs, are weights, is a bias, is the preactivation, and is the activation function. For a fully connected layer, this becomes , with a weight matrix . Most hidden-layer activations operate independently on each component, although some functions couple multiple components. (deeplearningbook.org)
In a multilayer perceptron, successive affine transformations without intervening nonlinearities collapse into a single affine transformation. For example,
Thus, merely adding such layers does not create nonlinear expressiveness. Nonlinear activations allow hidden representations to change the geometry of the input space, supporting representation learning rather than only repeated linear prediction. (en.d2l.ai)
Principal scalar functions
Identity and threshold functions. The identity activation, , leaves its input unchanged. It is useful where an unconstrained real-valued output is required. A threshold activation instead returns one of two values according to whether the input crosses a threshold. The classical perceptron uses a threshold decision, but the step function’s zero derivative away from its discontinuity makes ordinary gradient-based training unsuitable. (deeplearningbook.org)
Logistic sigmoid. The logistic function, commonly called the sigmoid in neural-network contexts, is
It maps finite real inputs into , and its derivative is . Its bounded output supports probability-valued predictions, as in logistic regression. For large positive or negative inputs, it saturates and its derivative approaches zero. (en.d2l.ai)
Hyperbolic tangent. The hyperbolic tangent activation maps inputs into , with derivative . Unlike sigmoid, its output range is symmetric around zero. It also saturates at both extremes, so its derivatives can become very small. (en.d2l.ai)
Rectified linear unit. The rectified linear unit (ReLU) is
It is zero for negative inputs and linear for positive inputs. Its derivative is zero on the negative side and one on the positive side; at zero, implementations adopt a convention because the ordinary derivative does not exist. A unit that remains on the negative side for all relevant inputs receives no gradient through its activation, a behavior often called a “dead ReLU.” (deeplearningbook.org)
Leaky and parametric rectifiers. Leaky ReLU replaces the negative branch with , usually with a small positive slope. Parametric ReLU (PReLU) learns that slope rather than fixing it. These variants preserve a nonzero derivative for negative inputs when . PReLU was introduced in a 2015 study that also developed initialization tailored to rectifier networks. (arxiv.org)
Smooth and exponential variants
The exponential linear unit (ELU) equals for nonnegative inputs and for negative inputs, where . It permits negative outputs while retaining a linear positive branch. Its negative branch saturates toward . (arxiv.org)
The Gaussian error linear unit (GELU), proposed in 2016, is
where is the cumulative distribution function of the standard normal distribution. Rather than applying ReLU’s hard sign-based cutoff, GELU continuously weights the input by . Implementations may use either its exact expression or an approximation. (arxiv.org)
The sigmoid linear unit (SiLU), also called Swish for this formulation, is . It is smooth and nonmonotonic, bounded below but unbounded above. Smoothness, saturation, output range, and computational cost distinguish these functions; nonlinearity does not require either monotonicity or boundedness. (keras.io)
Output activations and probability
Output activations have a different role from hidden-layer nonlinearities. For mutually exclusive classes, the softmax function transforms a vector of scores into positive components summing to one:
Unlike an elementwise activation, each component depends on every input score. A sigmoid can represent a binary-class probability, while separate sigmoids can represent multiple labels that need not be mutually exclusive. An identity output supports unconstrained regression, including models trained with mean squared error. (keras.io)
The activation and loss function must be interpreted together. Libraries may combine sigmoid and binary cross-entropy into one numerically stable operation receiving raw scores. In that arrangement, explicitly applying sigmoid before the loss would duplicate the transformation. (docs.pytorch.org)
Interaction with training
Backpropagation applies the chain rule through activation derivatives. Repeated multiplication by small derivatives can contribute to the vanishing gradient problem, especially in deep stacks of saturated units. Activation behavior also interacts with weight initialization: a 2010 study analyzed saturation and layerwise gradient propagation, while rectifier-specific initialization was developed subsequently. Activation choice therefore affects optimization alongside weight scales and network depth, rather than independently of them. (deeplearningbook.org)
Activations need not be fixed scalar formulas. PReLU contains learned parameters, and a gated linear unit splits an input into two parts and computes , where denotes elementwise multiplication. Software libraries consequently support both simple activation callables and dedicated activation layers that can maintain trainable state. (arxiv.org)