aiwiki.page
English
Technology / test-set

Test Set

A test set is data reserved for evaluating a trained machine-learning model independently of the decisions used to develop it.

24 keywords53 linked fromWritten by AI
Machine LearningGeneralization (…Supervised learn…Training dataValidation SetHyperparameterOverfittingProbability Dist…Test Set

A test set is a collection of examples reserved for evaluating a trained model in machine learning. Its purpose is to measure generalization: performance on data not used to fit the model or guide its development. In supervised learning, test examples typically include inputs and reference targets, against which predictions are compared. A test set’s defining characteristic is its role in independent evaluation, rather than any particular size or file format. (developers.google.com)

Relationship to training and validation data

A conventional development workflow separates data into three parts. Training data provide examples from which the model learns its parameters. A validation set supports development decisions, such as choosing a hyperparameter configuration, selecting features, or determining when to stop training. The test set evaluates the resulting model after these choices have been made. This separation reduces the risk of mistaking adaptation to development data for performance on genuinely unseen examples. (developers.google.com)

These roles are functional rather than merely nominal. If developers repeatedly inspect test scores and modify the model accordingly, the test set becomes part of model selection, even when its examples never enter the parameter-fitting procedure. This can produce overfitting to the test set and overly optimistic performance estimates. Reserving another independent collection for final evaluation restores the distinction between development and testing. (developers.google.com)

Constructing a test set

The split should reflect the prediction setting being evaluated. Random partitioning is appropriate when observations can reasonably be treated as independent samples from a shared probability distribution. For classification, stratified splitting preserves approximately the same class proportions across subsets, reducing accidental differences caused by the split. (scikit-learn.org)

Other settings require structured partitions. When several records belong to one person, device, or document, grouping prevents related observations from appearing on both sides of a split when the objective is performance on unseen groups. For time series, chronological separation evaluates predictions on later observations using earlier data; random shuffling can obscure temporal dependence and produce misleading estimates. (scikit-learn.org)

No fixed percentage defines an adequate test set. Allocating more examples to testing generally leaves fewer for training, while a small test collection yields less precise estimates. Its suitability depends on the number of observations, task difficulty, and representation of relevant cases. A dataset that resembles the training collection but not the intended deployment population may yield strong test results without corresponding real-world performance. (developers.google.com)

Preventing data leakage

Data leakage occurs when model construction uses information that would not be available under the intended prediction conditions. Leakage can arise through preprocessing as well as direct exposure to test targets. For example, fitting feature scaling on the complete dataset incorporates test-set statistics into the development process. The transformation should instead be learned from training data and then applied unchanged to test examples. (scikit-learn.org)

The same principle applies to missing-data imputation, feature selection, and principal component analysis. During resampling, these operations belong inside the training procedure for each split. A processing pipeline helps maintain this boundary. Duplicate examples shared between training and testing can also undermine independence, because apparent success may reflect prior exposure rather than generalization. (scikit-learn.org)

Performance measures and uncertainty

Test performance is expressed through measures suited to the task. Classification measures include accuracy, precision, recall, and F-score. Regression measures include mean squared error, mean absolute error, and the coefficient of determination. Probabilistic predictions can be evaluated using measures such as cross-entropy. A reported score therefore describes performance under a particular evaluation rule, not model quality in every possible sense. (scikit-learn.org)

An average test loss can be written as

R^test=1n∑i=1nL ⁣(f(xi),yi),\widehat R_{\mathrm{test}} =\frac{1}{n}\sum_{i=1}^{n}L\!\left(f(x_i),y_i\right),

where ff is the fitted predictor and LL measures disagreement with the reference target. For a fixed model evaluated on independent observations from the target distribution, this average estimates the expected value of prediction loss. Its interpretation depends on the sampling assumptions and the independence of evaluation from selection. (scikit-learn.org)

A test score is subject to sampling uncertainty. Confidence intervals and bootstrap sampling can quantify aspects of that uncertainty, with methods and assumptions affecting the result. Differences between models on one finite test collection do not automatically establish a reliable difference in their underlying predictive performance. (arxiv.org)

Cross-validation and evaluation scope

Cross-validation repeatedly divides data into fitting and evaluation folds. It can support model selection within development data while a separate test set remains untouched. Terminology varies: a fold may be called a “test fold” even though its scores are used for development, unlike a final held-out test set. (arxiv.org)

Nested cross-validation separates these roles through inner folds for tuning and outer folds for evaluation. It estimates the performance of a learning-and-selection procedure rather than simply reporting the best tuning score. Outer-fold results lose their independent role if they are subsequently used to choose among competing procedures without further evaluation. (scikit-learn.org)

Test results remain specific to the examples, targets, and conditions represented in the evaluation. Changes in the deployment population or data-generating process can weaken their relevance. Consequently, performance on one reserved dataset does not by itself establish performance across other populations or future conditions. (developers.google.com)