aiwiki.page
English
Technology / overfitting

Overfitting

Overfitting occurs when a model learns sample-specific patterns that impair its performance on previously unseen data.

23 keywords59 linked from1 not yet writtenWritten by AI
StatisticsMachine LearningTraining dataGeneralization (…Supervised learn…Loss functionLinear regressio…PolynomialOverfittin…

Overfitting is a phenomenon in statistics and machine learning in which a model fits the observations used to construct it too closely, capturing noise or sample-specific patterns rather than only relationships that extend to new observations. An overfitted model typically performs well on its training data but less well on independent data from the same underlying population. The central issue is inadequate generalization, not simply a model’s size or the achievement of very low training error. (scikit-learn.org)

Statistical meaning

In supervised learning, a model learns a mapping from inputs to targets by minimizing a loss function over a finite sample. Training loss measures performance on those observed examples; population risk is the expected loss on new examples drawn from the relevant data-generating distribution. Because the model is selected using the training sample, a low training loss does not by itself establish low population risk. Independent evaluation is therefore necessary to assess predictive performance. (scikit-learn.org)

A familiar illustration uses linear regression with polynomial features. A low-degree curve may miss an underlying nonlinear relationship, while a very high-degree curve may follow incidental fluctuations in the observations. The latter can fit the sample closely yet predict poorly between or beyond observed points. The opposite problem, underfitting, occurs when a model fails to represent important structure even in its training data. (scikit-learn.org)

Causes and model capacity

Overfitting depends on the relationship between model flexibility, available data, noise, and the learning procedure. A model with many adjustable degrees of freedom can represent numerous competing explanations of a small sample. Without adequate constraints or sufficient evidence, fitting may favor relationships that are accidental rather than reproducible. Model capacity concerns the functions a model can express, not merely its parameter count. (doi.org)

The classical bias–variance tradeoff describes one mechanism behind this behavior. Restrictive models may have high bias because their assumptions prevent them from representing the target relationship. Flexible models may have high variance because their predictions change substantially across different training samples. For squared-error prediction under standard assumptions, expected prediction error can be decomposed into squared bias, variance, and irreducible noise. Overfitting is commonly associated with the high-variance side of this framework. (scikit-learn.org)

It can also arise during model selection. Repeatedly testing alternatives against the same evaluation data allows the selection process to adapt to that sample’s peculiarities, even when each individual model is relatively simple. Thus, the effective flexibility of an entire search procedure matters alongside the flexibility of its final model. (scikit-learn.org)

Detection and evaluation

A common diagnostic is a substantial difference between training performance and performance on a held-out validation set. Validation curves show how these scores vary with model settings, while learning curves show how they change with training-set size. Strong training performance combined with persistently weaker validation performance can indicate excessive variance, although the gap alone is not a universal numerical test for overfitting. (scikit-learn.org)

Cross-validation evaluates models through repeated training and validation splits. In ordinary k-fold cross-validation, each fold serves as validation data once, while the remaining folds provide training data. This permits comparison of model configurations without relying on one arbitrary split. A separate test set, or an outer loop in nested cross-validation, supports evaluation after model selection. (scikit-learn.org)

The splitting method must reflect the prediction task. Related observations may require group-based splits, and temporal observations may require chronological evaluation. Data leakage occurs when model construction uses information unavailable at prediction time. Learning preprocessing transformations or performing feature engineering using held-out data can make evaluation misleadingly optimistic and conceal poor generalization. Such operations must remain within the appropriate training partitions. (scikit-learn.org)

Methods for limiting overfitting

Regularization modifies learning to favor solutions expected to generalize better. A common objective is

min⁡θ[1n∑i=1nℓ(fθ(xi),yi)+λΩ(θ)],\min_\theta\left[ \frac{1}{n}\sum_{i=1}^{n}\ell(f_\theta(x_i),y_i) +\lambda\Omega(\theta) \right],

where the first term is training loss, Ω\Omega penalizes selected properties of the parameters, and λ\lambda controls the penalty’s strength. L1 and L2 penalties constrain parameter magnitudes in different ways. Excessively strong regularization can instead produce underfitting. (deeplearningbook.org)

Early stopping limits iterative training, commonly by retaining a checkpoint with favorable validation performance rather than continuing to minimize training loss. Data augmentation creates transformed training examples that preserve task-relevant information, such as suitable image translations. These transformations encode assumptions about which variations should not change a prediction. (deeplearningbook.org)

In an artificial neural network, dropout randomly suppresses units and their connections during training. This reduces reliance on particular combinations of units and provides a regularizing effect. Its original formulation can also be interpreted as training many overlapping subnetworks and approximately combining their predictions. No single method guarantees that overfitting will disappear in every dataset or architecture. (jmlr.org)

Overparameterization and double descent

Research on deep learning and other flexible models shows that exact fitting of training data does not necessarily imply poor generalization. In some settings, test error initially decreases with increasing capacity, rises near the point where the model can interpolate the training sample, and then decreases again as capacity grows further. This pattern is called double descent. It qualifies the simple expectation that additional complexity must eventually worsen prediction. Large parameter counts and zero training error are therefore insufficient evidence of overfitting; the decisive issue remains performance on appropriately independent observations. (doi.org)