Cross-validation is a family of resampling procedures in statistics and machine learning used to estimate how well a predictive model performs on previously unseen observations. It repeatedly separates available data into fitting and evaluation subsets, trains a model on the former, and measures performance on the latter. Especially important in supervised learning, it supports both performance estimation and comparison of alternative modeling choices, while reducing dependence on a single data split. (scikit-learn.org)
Purpose and basic procedure
Performance measured on the same observations used for fitting can be misleading because a model may capture sample-specific patterns rather than relationships that support generalization. Cross-validation evaluates predictions for observations excluded from the corresponding fit, helping distinguish predictive performance from overfitting. Unlike a single holdout evaluation, it reuses observations across several fitting and evaluation rounds. (scikit-learn.org)
In ordinary k-fold cross-validation, a dataset is partitioned into disjoint, approximately equal-sized subsets called folds. For each fold, the learning procedure is fitted anew using the other folds as training data. The excluded fold serves as a validation set. Each observation is evaluated once, although it participates in training during the other rounds. Five-fold and ten-fold designs are common. (scikit-learn.org)
Let denote the fold index sets, and let be the predictor fitted without fold . For an observation-level loss function , an observation-weighted estimate is
For equal-sized folds, this equals the average of the fold-specific mean losses. Other evaluation measures can also be averaged across folds, but averaging a nonlinear metric need not equal computing it once from all held-out predictions. (arxiv.org)
Principal variants
Stratified k-fold cross-validation approximately preserves class proportions within each fold. This is useful when rare classes might otherwise be absent from evaluation subsets. Stratification addresses practical class-balance problems; it does not make observations independent or eliminate statistical uncertainty. (scikit-learn.org)
Leave-one-out cross-validation uses : each model is trained on all observations except one. It maximizes the training sample size within each round but generally requires fits. Special computational shortcuts exist for some models, so its cost depends on the learning procedure. (scikit-learn.org)
Repeated k-fold cross-validation repeats the partitioning process with different randomized folds. Repeated random-split validation, also called Monte Carlo cross-validation, instead repeatedly samples training and evaluation subsets of specified sizes. In the latter design, an observation may be evaluated several times or not at all. Repetition explores sensitivity to partitioning but does not create new independent observations. (scikit-learn.org)
Model selection and nested evaluation
Cross-validation can compare algorithms or tune hyperparameters, such as the strength of regularization in ridge regression or kernel settings in a support vector machine. However, selecting the configuration with the best cross-validation score and reporting that same score introduces selection bias: the search can exploit random fluctuations in the evaluation criterion. This is a form of overfitting at the model-selection level. (jmlr.org)
Nested cross-validation separates selection from assessment. Within each outer training subset, an inner cross-validation loop selects the configuration. That configuration is then fitted on the outer training subset and evaluated on the untouched outer fold. The outer results estimate the performance of the complete selection-and-fitting procedure, rather than merely its best inner score. (scikit-learn.org)
A separate test set offers another separation: cross-validation operates on development data, while final assessment uses observations excluded from model selection. After assessment, a model may be refitted using the selected settings and available development data. Cross-validation itself does not require deploying or averaging the individual fold models. (scikit-learn.org)
Dependence and data leakage
The splitting scheme must reflect the intended prediction task. Random observation-level splitting can be inappropriate when several records belong to the same person, organization, or other unit. Group-based splitting keeps such units out of both training and evaluation within a round when the objective is prediction for previously unseen groups. (scikit-learn.org)
For time series, chronological evaluation trains on earlier observations and evaluates later ones. Expanding-window or rolling-window designs preserve this direction; a gap can separate training from evaluation when required by the task. Randomly mixing past and future observations can produce an evaluation that does not represent forecasting conditions. (scikit-learn.org)
Data leakage also occurs when evaluation observations influence preprocessing. Learned scaling, missing-value imputation, feature engineering, feature selection, and principal component analysis must be fitted within the relevant training subset. A preprocessing pipeline can apply the resulting transformations to held-out observations without learning from them. Performing these operations on the complete dataset before splitting compromises the separation. (scikit-learn.org)
Interpretation and uncertainty
The choice of involves computational cost, training sample size, and the bias–variance tradeoff. Cross-validation evaluates models trained on subsets smaller than the full dataset, and its estimate is not automatically unbiased for the final full-data fit. Larger is therefore not universally superior. (arxiv.org)
Fold scores are dependent because their training sets overlap. Their standard deviation describes observed variation across folds, but is not automatically a standard error or valid confidence interval. Bengio and Grandvalet showed in 2004 that no universally unbiased estimator of the variance of k-fold cross-validation exists under all data distributions. Reporting the splitting design, evaluation measure, tuning procedure, and uncertainty assumptions is consequently essential to interpreting an estimate. (jmlr.org)