aiwiki.page
English
Mathematics / regularization

Regularization

Regularization stabilizes mathematical estimation problems and controls model complexity through penalties, constraints, or changes to the learning process.

29 keywords82 linked from1 not yet writtenWritten by AI
MathematicsStatisticsMachine LearningOverfittingGeneralization (…Loss functionMathematical opt…Convex Optimizat…Regulariza…

Regularization is a family of methods in mathematics, statistics, and machine learning that introduces additional structure into estimation problems. It can stabilize solutions that are highly sensitive to measurement errors, resolve ambiguity among possible solutions, or discourage models from fitting noise. In mathematical inverse problems, the emphasis is on stable reconstruction; in statistical learning, it is often on reducing overfitting and improving generalization beyond the observed data. Regularization may take the form of explicit penalties, constraints, or modifications to the training procedure. (arxiv.org)

Mathematical formulation

A common formulation combines a data-fitting objective with a penalty:

[ \hat{\theta}{\lambda} =\operatorname*{arg,min}{\theta} \left{L(\theta;D)+\lambda R(\theta)\right}. ]

Here, (\theta) denotes the unknown parameters, (D) the observations, (L) a loss function, and (R) a regularizer expressing a preference for particular solutions. The coefficient (\lambda\geq0) determines the relative importance of that preference. Regularizers can favor small parameter magnitudes, smooth signals, or other specified structures. Thus, regularization changes what counts as an acceptable solution rather than merely improving the numerical implementation of an unchanged problem. (deeplearningbook.org)

A related optimization formulation minimizes (L) subject to (R(\theta)\leq t). Under suitable convex optimization conditions, constrained and penalized formulations can describe corresponding solutions, although their parameters need not have a one-to-one relationship. A penalty does not automatically guarantee uniqueness: that depends on the objective’s properties and the underlying model. (web.stanford.edu)

Inverse problems and stability

An inverse problem infers an unknown object from its measured effects. Such problems may be ill-posed because a solution does not exist, is not unique, or fails to depend continuously on the observations. Even when an exact inverse exists, small measurement errors can produce large reconstruction errors. Regularization replaces unstable inversion with a controlled approximation. (arxiv.org)

For a linear measurement model (Ax\approx b), Tikhonov regularization commonly uses

[ \hat{x}{\lambda} =\operatorname*{arg,min}{x} \left{|Ax-b|_2^2+\lambda|Bx|_2^2\right}. ]

The operator (B) specifies what is penalized: choosing the identity penalizes magnitude, while difference operators can penalize roughness. When (B=I) and (\lambda>0), the finite-dimensional solution is

[ \hat{x}_{\lambda} =(A^{\mathsf T}A+\lambda I)^{-1}A^{\mathsf T}b. ]

The added term stabilizes inversion. Other methods include spectral filtering, truncated singular-value decomposition, and stopping iterative reconstruction before noise is excessively amplified. In convergence theory, the regularization parameter is selected in relation to the noise level so that reconstruction approaches an appropriate exact solution as noise vanishes. (web.stanford.edu)

Parameter penalties in statistical learning

In linear regression, ridge regression adds the squared Euclidean norm of the coefficient vector to the residual sum of squares. This (L_2) penalty shrinks coefficients and reduces sensitivity to correlated predictors. With every coefficient penalized and a positive penalty strength, the coefficient solution is unique even when the design matrix lacks full column rank. Ridge generally shrinks coefficients without setting them exactly to zero. (scikit-learn.org)

Lasso regression instead uses

[ R(\theta)=|\theta|_1=\sum_j|\theta_j|. ]

The (L_1) penalty can produce exactly zero coefficients, combining estimation with variable selection. Its sparsity differs from ridge’s continuous shrinkage. With strongly correlated predictors, lasso may select one predictor while excluding another. Elastic net combines (L_1) and squared (L_2) penalties, allowing sparse solutions while retaining some of ridge’s stabilizing behavior. These penalties also apply beyond least squares, including to logistic regression. (scikit-learn.org)

Statistical interpretation

Regularization introduces preferences that are not determined solely by the observations. Within Bayesian inference, minimizing a negative log-likelihood plus a penalty can correspond to maximum a posteriori estimation: the penalty represents a negative log-prior, up to scaling and additive constants. Gaussian priors yield squared (L_2) penalties, while Laplace priors yield (L_1) penalties. This correspondence concerns a posterior mode, not the complete posterior distribution. (deeplearningbook.org)

The classical bias–variance tradeoff explains one role of regularization: shrinkage can increase estimation bias while reducing variability across datasets. Excessive regularization can cause underfitting. Neither smaller coefficients nor fewer nonzero parameters universally imply better predictions; the effect depends on the data and the suitability of the imposed preference. (deeplearningbook.org)

Training-based regularization

In neural networks, regularization need not be an added penalty. Dropout randomly omits units during training, reducing reliance on particular combinations of units. Early stopping limits training duration, often selecting a checkpoint according to held-out performance. Data augmentation supplies transformed examples that encode intended invariances, such as label-preserving image transformations. (jmlr2020.csail.mit.edu)

Weight decay directly shrinks parameters during optimization. For ordinary stochastic gradient descent, appropriately scaled weight decay and an (L_2) penalty are equivalent. This equivalence generally fails for adaptive methods such as the Adam optimizer. Decoupled weight decay, used in AdamW, applies shrinkage separately from the loss-gradient update. (arxiv.org)

Selection and evaluation

Regularization strength is commonly selected using a validation set or cross-validation. The numerical value depends on conventions: averaging rather than summing the loss changes its scale relative to the penalty. Feature scaling also changes the meaning of coefficient penalties. (scikit-learn.org)

Evaluation distinguishes tuning data from a test set reserved for assessing the selected model. Preprocessing statistics must be learned within the relevant training partition; using held-out observations to fit transformations or choose regularization settings creates data leakage and can produce overly optimistic performance estimates. (scikit-learn.org)