Nested cross-validation is a model-evaluation procedure in machine learning and statistics that uses two levels of cross-validation to separate model selection from performance assessment. The inner level selects a model configuration, usually through hyperparameter optimization, while the outer level estimates the generalization performance of the complete selection-and-training procedure. Its central purpose is to avoid evaluating a selected configuration on the same validation results that led to its selection. (scikit-learn.org)
Motivation and selection bias
Ordinary cross-validation estimates performance by repeatedly fitting a model on one subset and evaluating it on another. However, when many configurations are compared, the configuration with the best validation score may benefit partly from random variation rather than genuinely superior predictive performance. Reporting that winning score as an independent performance estimate can therefore produce optimistic results. This is overfitting of the selection criterion, distinct from fitting model parameters too closely to the training observations. (jmlr.org)
Selection can concern a hyperparameter, such as regularization strength, or broader choices such as model family and feature selection. A study comparing support vector machines, random forests, and logistic regression may place all these alternatives inside the inner selection procedure. If their outer scores instead determine the winner, the reported winning outer score becomes subject to another layer of selection bias. Nesting separates selection from evaluation; it does not make unrestricted reuse of evaluation results harmless. (jmlr.org)
The two-level procedure
In a conventional implementation, the dataset is partitioned into outer folds. Each fold serves once as an outer test set, while the remaining observations form the outer training data. Within each outer training set, a separate -fold cross-validation creates inner training subsets and a rotating validation set. The outer test fold remains excluded throughout selection. (scikit-learn.org)
For each outer fold, the procedure is:
- Evaluate candidate configurations using only the inner folds.
- Select the configuration with the best aggregate inner score.
- Refit that configuration on the entire outer training set.
- Evaluate the refitted model once on the outer test fold.
The outer scores are then aggregated. Different outer folds can select different configurations; the object being evaluated is the selection procedure, not necessarily one fixed configuration. (scikit-learn.org)
Mathematical formulation
For supervised learning, let , and let contain the indices of outer test fold . Write for the remaining data. Let denote the full learning procedure, including inner selection and refitting, and define
For a pointwise loss function , a sample-weighted nested estimate is
This expresses the outer evaluation of the complete procedure. With equally sized folds, it equals the arithmetic mean of the fold-average losses. With unequal folds, weighting by fold size gives each observation equal weight. For nonlinear metrics, such as the F-score, averaging fold scores and scoring pooled predictions need not yield the same quantity. (scikit-learn.org)
The estimate concerns a procedure trained on the outer training-set size, approximately , rather than directly measuring a final model trained on all observations. Consequently, separation of tuning and testing does not guarantee exact absence of estimator bias for every target of interest. (sklearn.org)
Preprocessing and data leakage
Nesting is effective only when every data-dependent operation respects the split boundaries. Feature scaling, missing-data imputation, and principal component analysis must be fitted using the relevant training subset, then applied to validation or test observations without refitting on them. Selecting features using the complete dataset before cross-validation introduces data leakage, even if model fitting itself follows a nested design. (scikit-learn.org)
A processing pipeline keeps transformations and estimation together so that each inner training fold fits its own preprocessing steps. After selection, the entire pipeline is refitted on the outer training set. Outer test observations must not determine transformation parameters, feature choices, or other learned components. (scikit-learn.org)
The splitting strategy also determines the meaning of evaluation. Random folds suit settings where observations are sufficiently independent and exchangeable. Group-aware splits keep related observations together when performance on unseen groups is the target. For time series, chronological splits can evaluate prediction of later observations from earlier ones. These constraints apply at both nesting levels. (sklearn.org)
Computational cost and uncertainty
With outer folds, inner folds, and candidate configurations, exhaustive selection entails approximately candidate fits plus outer refits. This count follows directly from repeating the inner search for each outer fold. Parallel computing can reduce elapsed time, but does not remove the underlying fitting workload. (scikit-learn.org)
Outer scores show variation across partitions, but they are not independent replicates because their training sets overlap. Their standard deviation is therefore not automatically a valid standard error of the aggregate estimate. A confidence interval calculated as though fold scores were independent can understate uncertainty. Bengio and Grandvalet established that no universally unbiased variance estimator exists for ordinary -fold cross-validation. (jmlr.org)
Final fitting and reporting
After evaluation, the same predefined selection procedure can be run on the complete development dataset, followed by refitting the selected configuration on all available development observations. The nested scores remain estimates of the procedure’s performance, not fresh measurements of that final fitted model. A separately reserved test set provides another assessment only if it remains untouched by development decisions. (scikit-learn.org)
For reproducibility, an evaluation description records both splitting schemes, candidate configurations, scoring and aggregation rules, preprocessing steps, random seeds, and refitting behavior. These details specify which selection procedure the reported outer scores actually evaluate. (sklearn.org)