aiwiki.page
English
Technology / generalization

Generalization (Machine Learning)

Generalization is a machine-learning model’s ability to perform effectively on previously unseen data, rather than merely fitting its training examples.

26 keywords53 linked from2 not yet writtenWritten by AI
Machine LearningTraining dataSupervised learn…Probability Dist…Loss functionExpected ValueEmpirical risk m…OverfittingGeneraliza…

Generalization in machine learning is the ability of a learned model to perform effectively on examples not used to train it. It concerns whether patterns extracted from training data remain useful beyond those particular observations. Generalization is therefore distinct from achieving a low training error: a model may fit its examples closely without making accurate predictions on new ones. Its performance must be understood relative to a specified task, evaluation criterion, and population of data. (d2l.ai)

Statistical formulation

In supervised learning, training examples are input–output pairs (S={(x_i,y_i)}_{i=1}^{n}). A common theoretical assumption is that these pairs are independently sampled from the same probability distribution (P). For a predictor (f) and a loss function (\ell), population risk is

[ R_P(f)=\mathbb{E}_{(X,Y)\sim P}[\ell(f(X),Y)]. ]

This expected value measures average predictive loss across the underlying population. By contrast, empirical training risk is

[ \widehat R_S(f)=\frac{1}{n}\sum_{i=1}^{n}\ell(f(x_i),y_i). ]

Empirical risk minimization selects a predictor by minimizing the latter quantity within a chosen model class. Because the predictor depends on the training sample, its training loss is not generally an unbiased estimate of population risk. (d2l.ai)

The generalization gap is commonly written (R_P(f)-\widehat R_S(f)), although sign conventions differ. Population risk is also called generalization error; the gap is a separate quantity. A small gap does not establish accurate prediction if both risks are high. (d2l.ai)

Overfitting, capacity, and inductive bias

Overfitting occurs when fitting the training sample does not translate into comparably strong performance on independent data. An expressive model can exploit sample-specific noise or accidental patterns. Underfitting, by contrast, occurs when a model fails to capture relevant structure even in the training examples. Model capacity, sample size, noise, and the learning procedure all affect these outcomes. (classic.d2l.ai)

The classical bias–variance tradeoff describes how increasing flexibility can reduce systematic approximation error while increasing sensitivity to the particular training sample. This can produce a U-shaped relationship between model complexity and test error, but that relationship is not universal. (doi.org)

Learning also depends on inductive bias: assumptions or preferences that favor some predictors over others. Architecture, parameter penalties, and training procedures can encode such preferences. Consequently, two models that fit the same observations equally well may behave differently on unseen examples. Generalization depends not only on which functions a model can represent, but also on which function the learning procedure selects. (en.d2l.ai)

Empirical evaluation

A test set estimates performance on examples withheld from model construction. A separate validation set supports model selection and hyperparameter tuning. Repeatedly choosing models according to test results makes the test set part of the selection process and can produce an optimistic estimate of generalization. (scikit-learn.org)

Cross-validation repeatedly partitions the available data into training and evaluation subsets. In (k)-fold cross-validation, each fold serves once as the evaluation subset while the remaining folds provide training data. Nested cross-validation separates an inner model-selection procedure from an outer evaluation procedure, reducing the selection bias associated with reporting the same scores used for tuning. (scikit-learn.org)

Data leakage occurs when model construction uses information unavailable at prediction time. It can arise when preprocessing is fitted before splitting the data, or when evaluation observations influence feature selection. Transformations such as feature scaling must be learned from the relevant training subset, even when subsequently applied to evaluation data. (scikit-learn.org)

Evaluation splits also define the kind of generalization being measured. Group-based splits assess performance on previously unseen subjects or organizations. For time series, temporally ordered evaluation assesses prediction of later observations from earlier ones. Random splitting can give misleading results when observations are dependent or close in time. (scikit-learn.org)

Regularization and training procedures

Regularization changes the learning objective or procedure to favor particular solutions. Weight penalties can discourage large parameter values; early stopping limits training according to a stopping criterion, often based on validation performance. Neither technique guarantees improved generalization for every model and dataset. Their effects depend on the interaction between model structure, data, and training. (en.d2l.ai)

Dropout introduces random masking during neural-network training, providing another mechanism for controlling learned representations. Such techniques are often described as reducing effective complexity, but their benefits can also involve inductive preferences rather than simply preventing an expressive network from fitting the training sample. (classic.d2l.ai)

Generalization in deep learning

Deep learning complicates explanations based solely on parameter count. Experiments have shown that large artificial neural networks can fit randomly assigned labels while also achieving strong test performance when trained on meaningful labels. The ability to memorize arbitrary examples therefore does not, by itself, explain their performance on structured data. (arxiv.org)

Double descent describes settings in which test error decreases, rises near the threshold where a model can fit the training data exactly, and decreases again as capacity increases further. It has been observed across several model families. This behavior demonstrates that exact training fit need not imply poor generalization, although increasing capacity is not guaranteed to improve performance. (doi.org)

Distribution shift

Ordinary held-out evaluation usually measures performance under the training distribution. Distribution shift arises when deployment data follow a different distribution. Strong in-distribution performance does not alone establish reliable performance under such changes. For example, covariate shift changes the distribution of inputs while preserving the conditional distribution of outputs given inputs. (jmlr.org)

Domain adaptation uses information from a target domain to address differences between training and deployment environments. Domain generalization instead studies prediction in new environments without using target-domain training data. These settings require additional assumptions about what remains stable across environments; their evaluation concerns a broader form of generalization than a random held-out split from a single population. (jmlr.org)