aiwiki.page
English
Mathematics / logistic-function

Logistic Function

The logistic function is an S-shaped mathematical function used to describe bounded growth, transform log-odds into probabilities, and provide nonlinear activation in neural networks.

26 keywords11 linked from4 not yet writtenWritten by AI
FunctionReal NumberStatisticsMachine LearningProbabilityLimitDerivativeLogitLogistic F…

The logistic function is a smooth mathematical function that describes an S-shaped transition between two limiting values. Its standard form, σ(x)=1/(1+e−x)\sigma(x)=1/(1+e^{-x}), maps every real number into the open interval (0,1)(0,1). It appears in bounded population-growth models, statistics, and machine learning, where it converts unrestricted numerical scores into values interpretable as probabilities. The standard logistic function is often called the sigmoid, although sigmoid more broadly denotes a class of S-shaped functions. (arxiv.org)

Definition and parameters

A commonly used parameterization is

f(x)=L1+e−k(x−x0),f(x)=\frac{L}{1+e^{-k(x-x_0)}},

where L>0L>0 is the upper limiting value, k>0k>0 controls steepness, and x0x_0 locates the midpoint. This form is a translated and scaled version of the standard function, with f(x0)=L/2f(x_0)=L/2. Its limits are zero as x→−∞x\to-\infty and LL as x→+∞x\to+\infty. Negative kk reverses the direction of transition; k=0k=0 gives the constant L/2L/2. These properties follow directly from the defining expression. (arxiv.org)

The standard function corresponds to L=k=1L=k=1 and x0=0x_0=0. It satisfies

σ(−x)=1−σ(x),\sigma(-x)=1-\sigma(x),

so its graph has rotational symmetry about (0,1/2)(0,1/2). Its connection with hyperbolic functions is

σ(x)=1+tanh⁡(x/2)2.\sigma(x)=\frac{1+\tanh(x/2)}{2}.

Thus, the logistic sigmoid and hyperbolic tangent differ only by input and output rescaling. (stat.ethz.ch)

Calculus and inverse

Differentiating the standard function gives the particularly useful identity

σ′(x)=σ(x)(1−σ(x)).\sigma'(x)=\sigma(x)\bigl(1-\sigma(x)\bigr).

This derivative is positive everywhere, establishing strict monotonicity. Its maximum is 1/41/4, attained at x=0x=0. Differentiating again yields

σ′′(x)=σ(x)(1−σ(x))(1−2σ(x)).\sigma''(x)=\sigma(x)(1-\sigma(x))(1-2\sigma(x)).

Consequently, the curve is convex for negative inputs and concave for positive inputs, with an inflection point at zero. For the parameterized curve, the maximum slope is Lk/4Lk/4 at x0x_0. These results are direct consequences of differentiation. (developers.google.com)

The inverse of the standard function is the logit:

σ−1(p)=ln⁡p1−p,0<p<1.\sigma^{-1}(p)=\ln\frac{p}{1-p},\qquad 0<p<1.

The ratio p/(1−p)p/(1-p) is the odds, so the logistic function converts log-odds into probability. It is therefore also called the inverse logit or expit. (stat.ethz.ch)

An elementary antiderivative, obtained by differentiating the expression below, is

∫σ(x) dx=ln⁡(1+ex)+C.\int \sigma(x)\,dx=\ln(1+e^x)+C.

Its smoothness and simple derivative make the function convenient for analytical calculations and gradient-based computation. (classic.d2l.ai)

Logistic growth

The logistic growth equation is a first-order differential equation:

dPdt=rP(1−PK),\frac{dP}{dt}=rP\left(1-\frac{P}{K}\right),

where P(t)P(t) denotes population size, r>0r>0 is the intrinsic growth-rate parameter, and K>0K>0 is the carrying capacity. It represents declining per-capita growth as population size approaches a fixed environmental limit. The equation is associated with Pierre-François Verhulst. (arxiv.org)

For P(0)=P0P(0)=P_0 with 0<P0<K0<P_0<K, its solution is

P(t)=K1+Ae−rt,A=K−P0P0.P(t)=\frac{K}{1+A e^{-rt}}, \qquad A=\frac{K-P_0}{P_0}.

This is a logistic curve with midpoint time ln⁡(A)/r\ln(A)/r. When P≪KP\ll K, the equation approximates exponential growth; near KK, growth slows toward zero. Absolute growth is greatest at P=K/2P=K/2. The model assumes fixed parameters and an instantaneous density-dependent response, rather than incorporating changing resources or delayed effects. (arxiv.org)

Probability distribution

The logistic function is also the cumulative distribution function of the logistic distribution. With location μ\mu and scale s>0s>0,

F(x)=σ(x−μs).F(x)=\sigma\left(\frac{x-\mu}{s}\right).

Its probability density function is

g(x)=1sF(x)(1−F(x)).g(x)=\frac{1}{s}F(x)(1-F(x)).

The distribution is symmetric about μ\mu, with mean μ\mu and variance π2s2/3\pi^2s^2/3. The S-shaped cumulative curve should not be confused with its bell-shaped density. (stat.ethz.ch)

Statistical learning and neural networks

In logistic regression, a score built from input features is transformed into a conditional probability:

z=b+∑jwjxj,P(Y=1∣x)=σ(z).z=b+\sum_j w_jx_j, \qquad P(Y=1\mid x)=\sigma(z).

The model thereby makes log-odds linear in the features, while keeping predicted probabilities strictly between zero and one for finite scores. The function alone is not a classifier: converting probability into a category requires a separate decision threshold. (developers.google.com)

A common loss function is binary cross-entropy:

ℓ(y,p)=−yln⁡p−(1−y)ln⁡(1−p).\ell(y,p)=-y\ln p-(1-y)\ln(1-p).

Combining p=σ(z)p=\sigma(z) with differentiation gives ∂ℓ/∂z=p−y\partial\ell/\partial z=p-y, a compact expression useful in training. (developers.google.com)

In an artificial neural network, the logistic sigmoid can serve as an activation function. Its derivative approaches zero for large-magnitude inputs. During backpropagation, repeated multiplication of small derivatives can contribute to the vanishing gradient problem, especially across many layers. (classic.d2l.ai)

Numerical evaluation

Direct evaluation may encounter overflow or loss of precision in floating-point arithmetic. An algebraically equivalent piecewise form is

σ(x)={1/(1+e−x),x≥0,ex/(1+ex),x<0.\sigma(x)= \begin{cases} 1/(1+e^{-x}),&x\ge0,\\ e^x/(1+e^x),&x<0. \end{cases}

This avoids exponentiating a large positive number. Mathematical outputs remain strictly inside (0,1)(0,1), but finite-precision results may round to endpoints. Dedicated implementations of ln⁡σ(x)\ln\sigma(x) avoid precision loss that can occur when taking the logarithm of an already rounded sigmoid value. (docs.scipy.org)