The softmax function is a smooth mathematical function that transforms a finite vector of real-valued scores into a probability distribution over its components. Each output is proportional to the exponential of the corresponding input, and all outputs sum to one. In machine learning, softmax commonly serves as an output activation function for classifiers whose classes are mutually exclusive. It also normalizes weights inside models, rather than only producing final predictions. (docs.scipy.org)
Definition and interpretation
For (K) scores (z=(z_1,\ldots,z_K)), with each (z_i) a finite real number, softmax is defined by
[ p_i=\operatorname{softmax}(z)i =\frac{\exp(z_i)}{\sum{j=1}^{K}\exp(z_j)}, \qquad i=1,\ldots,K. ]
Thus (p_i>0) and (\sum_i p_i=1). When (K>1), every component is strictly less than one. These properties follow directly from the positive exponential terms in the definition. (docs.scipy.org)
In classification, the inputs are often called logits, meaning unnormalized class scores. They need not be positive or sum to one. This usage is related to, but broader than, the binary logit defined as log-odds. For example, direct substitution of (z=(0,\log 2,\log 3)) gives (p=(1/6,1/3,1/2)). The output assigns the largest probability to the largest score without discarding the other alternatives. (proceedings.mlr.press)
For two components, algebraic simplification gives
[ p_1=\frac{1}{1+\exp(-(z_1-z_2))}. ]
Consequently, binary softmax reduces to the logistic function applied to a score difference, connecting it with binary logistic regression. (deeplearningbook.org)
Mathematical properties
Softmax is invariant under adding a common constant:
[ \operatorname{softmax}(z+c\mathbf 1) =\operatorname{softmax}(z). ]
Only relative scores matter. In particular, taking the ratio of two components yields
[ \frac{p_i}{p_j}=\exp(z_i-z_j), \qquad \log\frac{p_i}{p_j}=z_i-z_j. ]
These identities, obtained directly from the definition, explain both its preservation of score ordering and its redundant common-offset parameterization. (deeplearningbook.org)
The Jacobian matrix describes how every output changes with every input. Differentiating the definition gives the partial derivatives
[ \frac{\partial p_i}{\partial z_j} =p_i(\delta_{ij}-p_j), \qquad J=\operatorname{diag}(p)-pp^\mathsf T, ]
where (\delta_{ij}) equals one when (i=j) and zero otherwise. Thus increasing one score increases its own probability and decreases the others. This coupling distinguishes softmax from independently applied scalar activations. (docs.scipy.org)
Softmax is also the gradient of the log-sum-exp function,
[ L(z)=\log\sum_j \exp(z_j). ]
Despite its name, softmax returns a vector, not an approximation to the scalar maximum. It can instead be understood as a smooth counterpart of selecting a maximizing index, expressed using one-hot encoding. (docs.scipy.org)
Classification and training
In multinomial logistic regression, softmax normalizes class scores produced by linear predictors. In an artificial neural network, the scores can instead depend on learned nonlinear representations. Both constructions specify a conditional distribution over class labels. (deeplearningbook.org)
For a target distribution (y), the cross-entropy loss is
[ \mathcal L=-\sum_i y_i\log p_i, \qquad \sum_i y_i=1. ]
Combining the preceding derivative with this loss gives
[ \frac{\partial\mathcal L}{\partial z_i}=p_i-y_i. ]
For a one-hot target with correct class (c), the loss reduces to (-\log p_c). Minimizing it corresponds to maximum likelihood estimation of the class distribution; the resulting derivatives propagate through earlier layers using backpropagation. (deeplearningbook.org)
Temperature and calibration
A positive temperature parameter (T) modifies the transformation:
[ p_i(T)=\frac{\exp(z_i/T)} {\sum_j\exp(z_j/T)}. ]
For fixed scores, larger temperatures flatten the distribution, while smaller temperatures concentrate it on high-scoring components. Directly from this formula, as (T\to\infty), the distribution approaches uniformity. As (T\to0^+), it concentrates on the maximum; if several scores share that maximum, their limiting probabilities are equal. Positive temperature scaling preserves score ordering. (proceedings.mlr.press)
A normalized output is not automatically well calibrated: predicted confidence need not match observed correctness frequencies. Temperature scaling fits a single positive parameter on a held-out validation set, typically by minimizing negative log-likelihood. Research has demonstrated improved calibration across several neural-network classification datasets without changing the highest-ranked class. (proceedings.mlr.press)
Numerical computation and attention
Direct exponentiation can overflow for large scores. A mathematically equivalent stable formulation subtracts (m=\max_j z_j):
[ p_i=\frac{\exp(z_i-m)} {\sum_j\exp(z_j-m)}. ]
All exponential arguments are then nonpositive, and at least one exponential equals one. Very small terms may still underflow. Numerical analysis supports this shifted formulation as an accurate way to avoid overflow and reduce harmful underflow. Software must also specify the axis along which normalization occurs. (academic.oup.com)
When logarithmic probabilities are needed, computing log-softmax directly avoids taking the logarithm of a probability already rounded to zero:
[ \log p_i=z_i-m-\log\sum_j\exp(z_j-m). ]
Cross-entropy implementations therefore commonly combine normalization and loss computation directly from logits. (deeplearningbook.org)
Within the Transformer architecture, an attention mechanism applies softmax to scaled query–key similarities:
[ \operatorname{Attention}(Q,K,V) =\operatorname{softmax}\left(\frac{QK^\mathsf T}{\sqrt{d_k}}\right)V. ]
Normalization is performed across keys for each query. The resulting weights determine a weighted combination of value vectors; here softmax allocates attention rather than class probabilities. Masks exclude disallowed positions by assigning their scores negative infinity before normalization. (arxiv.org)