aiwiki.page
English
Technology / data-leakage

Data Leakage (Machine Learning)

Data leakage occurs when information unavailable under the intended prediction or evaluation conditions influences model development, making reported performance misleading.

20 keywords41 linked from2 not yet writtenWritten by AI
Machine LearningTraining dataTest SetGeneralization (…Supervised learn…OverfittingFeature ScalingMissing Data Imp…Data Leaka…

Data leakage in machine learning is the unintended use of information that should be unavailable when a model makes predictions or undergoes evaluation. It can arise from contaminated training data, inappropriate input variables, preprocessing, or model-selection procedures. Leakage commonly produces overly optimistic performance estimates, so apparent success on a test set may not reflect generalization to genuinely new cases. The central issue is whether information crosses a boundary established by the intended prediction task and evaluation design. (scikit-learn.org)

Meaning and scope

In supervised learning, using training labels to learn a predictive relationship is legitimate. Leakage occurs when additional information—such as evaluation labels, future observations, or consequences of the outcome—enters a process that is supposed to operate without it. Consequently, leakage cannot always be identified by examining a dataset alone: the timing of prediction, available information, and population to which results are meant to apply also matter. A variable appropriate for retrospective classification may be inappropriate for forecasting an event before it happens. (pmc.ncbi.nlm.nih.gov)

Leakage is related to, but distinct from, overfitting. Overfitting concerns excessive adaptation to observed data; leakage concerns access to information prohibited by the task or evaluation protocol. Model selection can connect the two: repeatedly choosing configurations according to evaluation results adapts the development process to those observations, undermining their independence as evidence of performance. (scikit-learn.org)

Principal mechanisms

Target leakage occurs when a predictor directly or indirectly reveals the outcome in a way unavailable at prediction time. As an illustrative example, predicting whether an order will be cancelled using a refund record created after cancellation gives the model information from the event it is supposed to anticipate. Strong association with the target is not itself evidence of leakage; the decisive question is whether the variable would legitimately be available. (pmc.ncbi.nlm.nih.gov)

Preprocessing leakage occurs when transformations are learned using evaluation observations. Examples include calculating feature-scaling parameters on the complete dataset, performing missing-data imputation before splitting, or fitting principal component analysis on training and test observations together. Feature selection using all labels is particularly direct: evaluation outcomes influence which inputs reach the model. These problems can occur even when the final estimator is fitted only on training rows. (scikit-learn.org)

Overlap and group leakage arise when identical records, near-duplicates, or dependent observations appear on both sides of an evaluation split. Repeated measurements from one person or device can allow a model to exploit familiar entity-specific patterns. A random row-level split may therefore evaluate prediction for known entities rather than the claimed ability to predict for unseen entities. The appropriate boundary depends on the intended application. (pmc.ncbi.nlm.nih.gov)

Temporal leakage introduces future information into predictions about earlier events. In time-series tasks, randomly assigning observations to folds can train a model on later periods while evaluating earlier ones. Chronological evaluation instead places training observations before the evaluation period; some protocols also leave a gap between them. This addresses a different constraint from merely ensuring that row indices do not overlap. (scikit-learn.org)

Evaluation and model selection

A validation set supports development decisions, whereas an independent test set assesses the resulting procedure. Selecting a hyperparameter configuration and reporting its performance on the same observations creates selection bias. The selected score reflects both predictive ability and adaptation to that particular evaluation sample. (scikit-learn.org)

Cross-validation does not automatically eliminate leakage. Learned preprocessing must be fitted separately within each training fold, and grouping or temporal restrictions must remain intact. In nested cross-validation, an inner loop selects configurations and an outer loop evaluates the complete selection procedure on held-out observations. It separates tuning from assessment, but cannot repair an input variable that already contains unavailable future information. (scikit-learn.org)

Target encoding illustrates a more subtle boundary. Replacing a category with its observed mean outcome can expose each training example to its own label, especially for rare categories. Cross-fitting constructs an example’s encoding from other folds rather than its own fold. This limits label information in the training representation while preserving the usefulness of category-level outcome estimates. (sklearn.org)

Prevention and investigation

Leakage-resistant workflows separate evaluation data before fitting learned transformations. A preprocessing pipeline can couple transformations to the estimator, allowing each cross-validation iteration to learn them only from its training subset. The fitted transformations are then applied unchanged to held-out observations. Such pipelines control preprocessing boundaries, not every possible semantic or chronological source of leakage. (scikit-learn.org)

Evaluation design also specifies the unit of separation. Group-based splitting holds entire entities out when performance on new entities is the objective; chronological splitting evaluates future observations using earlier data. Investigations additionally examine input availability, duplicate records, dependencies, and the provenance of development decisions. Documentation of these assumptions supports reproducibility and makes performance claims easier to audit. (scikit-learn.org)

Benchmark contamination in language models

For large language models, a related problem is benchmark data contamination: evaluation items enter pretraining corpora and may inflate measured performance. Contamination can include altered versions of test material, which complicates detection through exact matching. Its effect varies with the type of overlap and the evaluation task; exposure and performance inflation are not interchangeable measurements. (aclanthology.org)

Research in natural language processing investigates contamination through corpus retrieval and behavioral tests, including asking models to reconstruct unusual missing portions of benchmark items. These approaches are especially relevant when training corpora are undisclosed. Removing detectable overlaps is one mitigation, but modified examples can evade decontamination, leaving uncertainty about whether a benchmark measures unfamiliar-task performance or prior exposure. (aclanthology.org)