Adam, short for adaptive moment estimation, is an algorithm for mathematical optimization that updates parameters using moving averages of gradients and squared gradients. It is used in machine learning, particularly to train artificial neural networks in deep learning. Adam combines a momentum-like estimate of update direction with adaptive scaling for each parameter. Diederik P. Kingma and Jimmy Ba introduced it in a preprint submitted on December 22, 2014; the paper appeared at ICLR 2015. (arxiv.org)
Optimization setting and origins
Training commonly involves minimizing an objective function built from a loss function evaluated on training data. Rather than compute an exact gradient over the entire dataset at every iteration, an optimizer can use gradients estimated from individual examples or minibatches. Adam operates in this stochastic setting and requires first-order derivatives, not second-order derivative matrices. (arxiv.org)
Adam draws on two developments in gradient descent: momentum, which accumulates information about past gradient directions, and coordinate-wise adaptive methods such as AdaGrad and RMSProp. AdaGrad accumulates squared gradients throughout training, whereas RMSProp and Adam use exponentially weighted histories. This reduces the persistent shrinkage of effective learning rates that can occur when squared gradients accumulate without forgetting. (sanjivk.com)
Update equations
Let (\theta_{t-1}) denote the parameter vector before iteration (t), and let
[ g_t=\nabla_\theta f_t(\theta_{t-1}) ]
be the gradient of the current stochastic objective. Adam initializes two state vectors, (m_0=0) and (v_0=0), and updates them as
[ m_t=\beta_1m_{t-1}+(1-\beta_1)g_t, ]
[ v_t=\beta_2v_{t-1}+(1-\beta_2)g_t^2. ]
Here (g_t^2) is an element-wise square, and (0\leq\beta_1,\beta_2<1) control how rapidly previous information is forgotten. The first vector smooths gradient direction; the second smooths squared gradient magnitude. (docs.pytorch.org)
Because both vectors start at zero, their early values are biased toward zero. Adam applies bias corrections:
[ \widehat m_t=\frac{m_t}{1-\beta_1^t}, \qquad \widehat v_t=\frac{v_t}{1-\beta_2^t}. ]
The standard parameter update is
[ \theta_t=\theta_{t-1} -\alpha_t \frac{\widehat m_t}{\sqrt{\widehat v_t}+\epsilon}. ]
All division and square-root operations are element-wise. The scalar (\alpha_t) is the base learning rate, while (\epsilon>0) stabilizes the denominator. Thus, parameters share a base step size but receive different adaptive scaling. (docs.pytorch.org)
Statistical interpretation
The moving averages estimate the first moment and the uncentered second moment of gradients. The latter is not the variance: variance subtracts the squared mean from the second moment. Adam therefore does not directly divide its update by an estimated gradient standard deviation. Instead, it normalizes a smoothed gradient by a root-mean-square-like measure of its recent magnitude. (arxiv.org)
Bias correction compensates for initialization rather than guaranteeing that a changing gradient process is estimated without error. For a stationary mean gradient, the expected uncorrected first-moment estimate contains the factor (1-\beta_1^t), which the correction removes. In neural-network training, the gradient distribution generally changes as parameters change, so the corrected quantities remain adaptive summaries of recent history. (arxiv.org)
Hyperparameters and implementation
Common hyperparameter values are (\alpha=0.001), (\beta_1=0.9), and (\beta_2=0.999). Larger decay coefficients retain past information longer. Adaptive scaling does not eliminate the base learning rate: implementations can use a constant value or a learning-rate schedule. These values are conventional defaults, not guarantees of optimal behavior for every objective. (tensorflow.org)
The stabilizing constant supports numerical stability, but implementations differ. PyTorch documents a default (\epsilon=10^{-8}); TensorFlow documents (10^{-7}) and identifies its epsilon with a rearranged formulation in the original paper rather than directly with Algorithm 1. Such differences matter when comparing nominally identical optimizer configurations. (docs.pytorch.org)
Adam stores two persistent state arrays matching the trainable parameters. Its state storage and element-wise update work therefore grow linearly with parameter count, excluding gradient computation. Optimizer state is distinct from model weights: saving only weights does not preserve the accumulated moment estimates. TensorFlow exposes separate methods for saving and loading optimizer variables. (tensorflow.org)
Convergence and limitations
Practical success does not establish universal convergence. In On the Convergence of Adam and Beyond, published at ICLR 2018, Sashank Reddi, Satyen Kale, and Sanjiv Kumar identified a flaw in the original convergence analysis and constructed convex optimization problems where Adam fails to reach the optimum. Exponential averaging can forget infrequent but important gradients, producing unfavorable changes in effective step sizes. (research.google)
Their AMSGrad variant retains the coordinate-wise maximum of past second-moment accumulators. This supplies longer-term memory and supports convergence guarantees under specified assumptions and step-size schedules. Those guarantees concern the analyzed optimization setting; they do not imply that every neural-network training run reaches a global minimum. (sanjivk.com)
Weight decay and AdamW
Regularization introduces another distinction. Adding an (L_2) penalty to the loss contributes a term proportional to the parameters to the gradient. In ordinary Adam, that contribution enters both moment accumulators and is adaptively rescaled. It is therefore not equivalent to independently shrinking the parameters, although the two procedures are equivalent for standard stochastic gradient descent after appropriate learning-rate rescaling. (arxiv.org)
AdamW, developed by Ilya Loshchilov and Frank Hutter, decouples weight decay from the gradient-based update. In a common convention,
[ \theta_t=(1-\alpha_t\lambda)\theta_{t-1} -\alpha_t\frac{\widehat m_t}{\sqrt{\widehat v_t}+\epsilon}. ]
The shrinkage term does not enter the moment estimates. The authors reported improved generalization on their image-classification experiments. AdamW changes how regularization is applied; AMSGrad instead changes the adaptive denominator to address convergence behavior. (arxiv.org)