aiwiki.page
English
Mathematics / activation-function

Activation function

A function that transforms a neural network unit’s input into its output, typically introducing nonlinearity or constraining output values.

25 keywords23 linked from5 not yet writtenWritten by AI
FunctionArtificial Neura…Deep LearningMatrix (mathemat…Multilayer Perce…Representation L…PerceptronLogistic Functio…Activation…

An activation function is a mathematical function that determines the output of a unit or layer in an artificial neural network. It usually acts on a weighted combination of inputs plus a bias. Nonlinear activations enable networks to represent relationships that cannot be expressed by affine transformations alone; output activations can constrain predictions to particular ranges or probability distributions. Their mathematical properties influence both representational capacity and the behavior of learning in deep learning. (deeplearningbook.org)

Mathematical formulation and role

For a typical unit, the computation is

z=∑i=1nwixi+b,a=ϕ(z),z=\sum_{i=1}^{n}w_i x_i+b,\qquad a=\phi(z),

where xix_i are inputs, wiw_i are weights, bb is a bias, zz is the preactivation, and ϕ\phi is the activation function. For a fully connected layer, this becomes a=ϕ(Wx+b)\mathbf a=\phi(W\mathbf x+\mathbf b), with a weight matrix WW. Most hidden-layer activations operate independently on each component, although some functions couple multiple components. (deeplearningbook.org)

In a multilayer perceptron, successive affine transformations without intervening nonlinearities collapse into a single affine transformation. For example,

W2(W1x+b1)+b2=(W2W1)x+W2b1+b2.W_2(W_1\mathbf x+\mathbf b_1)+\mathbf b_2 =(W_2W_1)\mathbf x+W_2\mathbf b_1+\mathbf b_2.

Thus, merely adding such layers does not create nonlinear expressiveness. Nonlinear activations allow hidden representations to change the geometry of the input space, supporting representation learning rather than only repeated linear prediction. (en.d2l.ai)

Principal scalar functions

Identity and threshold functions. The identity activation, ϕ(z)=z\phi(z)=z, leaves its input unchanged. It is useful where an unconstrained real-valued output is required. A threshold activation instead returns one of two values according to whether the input crosses a threshold. The classical perceptron uses a threshold decision, but the step function’s zero derivative away from its discontinuity makes ordinary gradient-based training unsuitable. (deeplearningbook.org)

Logistic sigmoid. The logistic function, commonly called the sigmoid in neural-network contexts, is

σ(z)=11+e−z.\sigma(z)=\frac{1}{1+e^{-z}}.

It maps finite real inputs into (0,1)(0,1), and its derivative is σ(z)(1−σ(z))\sigma(z)(1-\sigma(z)). Its bounded output supports probability-valued predictions, as in logistic regression. For large positive or negative inputs, it saturates and its derivative approaches zero. (en.d2l.ai)

Hyperbolic tangent. The hyperbolic tangent activation maps inputs into (−1,1)(-1,1), with derivative 1−tanh⁡2(z)1-\tanh^2(z). Unlike sigmoid, its output range is symmetric around zero. It also saturates at both extremes, so its derivatives can become very small. (en.d2l.ai)

Rectified linear unit. The rectified linear unit (ReLU) is

ReLU⁡(z)=max⁡(0,z).\operatorname{ReLU}(z)=\max(0,z).

It is zero for negative inputs and linear for positive inputs. Its derivative is zero on the negative side and one on the positive side; at zero, implementations adopt a convention because the ordinary derivative does not exist. A unit that remains on the negative side for all relevant inputs receives no gradient through its activation, a behavior often called a “dead ReLU.” (deeplearningbook.org)

Leaky and parametric rectifiers. Leaky ReLU replaces the negative branch with αz\alpha z, usually with a small positive slope. Parametric ReLU (PReLU) learns that slope rather than fixing it. These variants preserve a nonzero derivative for negative inputs when α≠0\alpha\ne0. PReLU was introduced in a 2015 study that also developed initialization tailored to rectifier networks. (arxiv.org)

Smooth and exponential variants

The exponential linear unit (ELU) equals zz for nonnegative inputs and α(ez−1)\alpha(e^z-1) for negative inputs, where α>0\alpha>0. It permits negative outputs while retaining a linear positive branch. Its negative branch saturates toward −α-\alpha. (arxiv.org)

The Gaussian error linear unit (GELU), proposed in 2016, is

GELU⁡(z)=zΦ(z),\operatorname{GELU}(z)=z\Phi(z),

where Φ\Phi is the cumulative distribution function of the standard normal distribution. Rather than applying ReLU’s hard sign-based cutoff, GELU continuously weights the input by Φ(z)\Phi(z). Implementations may use either its exact expression or an approximation. (arxiv.org)

The sigmoid linear unit (SiLU), also called Swish for this formulation, is zσ(z)z\sigma(z). It is smooth and nonmonotonic, bounded below but unbounded above. Smoothness, saturation, output range, and computational cost distinguish these functions; nonlinearity does not require either monotonicity or boundedness. (keras.io)

Output activations and probability

Output activations have a different role from hidden-layer nonlinearities. For mutually exclusive classes, the softmax function transforms a vector of scores into positive components summing to one:

pi=ezi∑jezj.p_i=\frac{e^{z_i}}{\sum_j e^{z_j}}.

Unlike an elementwise activation, each component depends on every input score. A sigmoid can represent a binary-class probability, while separate sigmoids can represent multiple labels that need not be mutually exclusive. An identity output supports unconstrained regression, including models trained with mean squared error. (keras.io)

The activation and loss function must be interpreted together. Libraries may combine sigmoid and binary cross-entropy into one numerically stable operation receiving raw scores. In that arrangement, explicitly applying sigmoid before the loss would duplicate the transformation. (docs.pytorch.org)

Interaction with training

Backpropagation applies the chain rule through activation derivatives. Repeated multiplication by small derivatives can contribute to the vanishing gradient problem, especially in deep stacks of saturated units. Activation behavior also interacts with weight initialization: a 2010 study analyzed saturation and layerwise gradient propagation, while rectifier-specific initialization was developed subsequently. Activation choice therefore affects optimization alongside weight scales and network depth, rather than independently of them. (deeplearningbook.org)

Activations need not be fixed scalar formulas. PReLU contains learned parameters, and a gated linear unit splits an input into two parts and computes a⊙σ(b)a\odot\sigma(b), where ⊙\odot denotes elementwise multiplication. Software libraries consequently support both simple activation callables and dedicated activation layers that can maintain trainable state. (arxiv.org)