Ridge regression is a method of linear regression that adds a penalty on the squared magnitude of coefficients to the least-squares fitting criterion. This form of regularization stabilizes estimation when predictors contain overlapping information and can improve prediction by accepting some bias in exchange for lower variability. It is used in statistics and machine learning, particularly when unpenalized estimates are sensitive to small changes in the observations. (scikit-learn.org)
Historical background and motivation
Arthur E. Hoerl and Robert W. Kennard developed an influential statistical formulation in their February 1970 paper, “Ridge Regression: Biased Estimation for Nonorthogonal Problems,” published in Technometrics. Their approach modifies the least-squares equations by adding positive quantities to the diagonal of the predictor cross-product matrix. They also introduced the ridge trace, a plot showing how estimated coefficients change with the penalty parameter. (homepages.math.uic.edu)
A central motivation is multicollinearity: predictors may be approximately expressible as linear combinations of other predictors. In such settings, ordinary least squares can produce coefficients with large variance, even when fitted values appear reasonable. Ridge regression reduces sensitivity to these weakly identified coefficient directions rather than removing predictors. (scikit-learn.org)
Mathematical formulation
Let be an predictor matrix, an -dimensional response vector, and the coefficient vector. Assuming that predictors and responses have been centered, ridge regression solves
The first term is the squared-error loss function; the second is the penalty, with . The hyperparameter controls their relative importance. Some formulations divide the residual sum of squares by , so numerical penalty values are not directly comparable without checking the convention. (scikit-learn.org)
The solution is
where is the identity matrix. For , the coefficient solution is unique even when lacks full column rank, including when there are more predictors than observations. An intercept is normally estimated separately and left unpenalized; centering permits the displayed formulation without an explicit intercept. (arxiv.org)
At , the criterion becomes ordinary least squares. As grows, the coefficient vector’s norm decreases; as , the penalized coefficients approach zero. Individual coefficients need not decrease monotonically in absolute value. (homepages.math.uic.edu)
Shrinkage and statistical interpretation
Ridge regression illustrates the bias–variance tradeoff. Under a correctly specified linear model with zero-mean errors, its coefficient estimates are generally biased toward zero. The variance reduction can nevertheless outweigh the squared bias, yielding lower mean squared error than least squares for an appropriate penalty. Improvement is not guaranteed for every dataset or penalty choice. (homepages.math.uic.edu)
The singular value decomposition provides a precise account of shrinkage. If , with singular values , the fitted response is
Directions with small singular values receive stronger attenuation. These are the directions in which unregularized coefficient estimation is most unstable. Thus, ridge does not simply multiply every original coefficient by one common factor. (arxiv.org)
In Bayesian inference, independent zero-centered Gaussian coefficient priors provide another interpretation. With Gaussian observation errors of variance and coefficient prior variance , maximum a posteriori estimation gives the stated ridge objective with . (arxiv.org)
Scaling, tuning, and computation
Ridge depends on predictor units because its penalty acts directly on coefficient magnitudes. Feature scaling, often to unit variance, makes penalties comparable across predictors measured on different scales. This is a modeling choice rather than a requirement that predictors follow a normal distribution. (scikit-learn.org)
The penalty is commonly selected through cross-validation, comparing prediction errors across candidate values. The purpose is to estimate generalization performance rather than minimize training residuals alone. A held-out test set is reserved for evaluation rather than parameter selection. (scikit-learn.org)
Scaling parameters must be learned from the training data within each validation split. Computing them from all observations before validation introduces data leakage, potentially producing optimistic performance estimates. Pipelines combine preprocessing and estimation so that each operation is fitted on the appropriate subset. (scikit-learn.org)
Although the closed-form solution contains an inverse, numerical implementations can solve the corresponding linear system instead. Available approaches include singular-value decomposition, Cholesky factorization, and iterative solvers; their suitability depends on matrix size, sparsity, and conditioning. (scikit-learn.org)
Related methods and limitations
Ridge is a special case of Tikhonov regularization, whose more general penalties can constrain selected combinations of coefficients. Unlike lasso regression, ridge generally does not produce exact zero coefficients and therefore does not inherently perform feature selection. Elastic net combines squared-magnitude and absolute-value penalties, incorporating both shrinkage and sparsity. (scikit-learn.org)
Through a kernel method, ridge can also be extended to nonlinear prediction while retaining a regularized least-squares objective. Ordinary ridge itself remains linear in its supplied features: shrinkage does not automatically correct an unsuitable functional form, and correlated predictors can still make individual coefficients difficult to interpret. (scikit-learn.org)