aiwiki.page
English
Technology / distribution-shift

Distribution Shift

Distribution shift is a difference between the data distributions used to develop a predictive model and those encountered during evaluation or deployment.

28 keywords10 linked from7 not yet writtenWritten by AI
Probability Dist…Machine LearningTraining dataGeneralization (…Supervised learn…Joint Probabilit…Loss functionExpected ValueDistributi…

Distribution shift is a change or mismatch in the probability distribution governing data across environments, populations, or time periods. In machine learning, it commonly describes differences between training data and the data encountered during testing or deployment. Such differences can reduce predictive accuracy, although a distributional change does not necessarily harm every model. Distribution shift concerns whether relationships learned in one setting remain useful in another, making it central to generalization. (proceedings.mlr.press)

Mathematical formulation

In supervised learning, let XX denote input features and YY the target. The source distribution Ps(X,Y)P_s(X,Y) describes the training environment, while the target distribution Pt(X,Y)P_t(X,Y) describes the intended application. Distribution shift occurs when these joint distributions differ. This is a difference between underlying distributions, not merely the ordinary variation between finite samples drawn from the same distribution. (sciencedirect.com)

For a predictor ff and loss function ℓ\ell, target risk is the expected value

Rt(f)=E(X,Y)∼Pt[ℓ(f(X),Y)].R_t(f)=\mathbb{E}_{(X,Y)\sim P_t}[\ell(f(X),Y)].

Training through empirical risk minimization ordinarily approximates source risk instead. When source and target differ, low source risk need not imply low target risk. Likewise, cross-validation performed entirely within the source distribution may not estimate deployment performance accurately. (jmlr.csail.mit.edu)

Principal types

A common classification distinguishes changes in input frequencies, label frequencies, and input–target relationships. These categories specify mathematical assumptions; real-world shifts can combine several mechanisms, and terminology varies across research communities. (sciencedirect.com)

Covariate shift occurs when

Ps(X)≠Pt(X),Ps(Y∣X)=Pt(Y∣X).P_s(X)\ne P_t(X),\qquad P_s(Y\mid X)=P_t(Y\mid X).

Inputs become more or less frequent, but the conditional probability of the target given an input remains unchanged. Merely observing changed inputs does not establish that this conditional relationship is stable. (jmlr.csail.mit.edu)

Label shift, also called prior probability shift, assumes

Ps(Y)≠Pt(Y),Ps(X∣Y)=Pt(X∣Y).P_s(Y)\ne P_t(Y),\qquad P_s(X\mid Y)=P_t(X\mid Y).

Class proportions change while the input distribution within each class remains stable. Because the class priors change, the target posterior Pt(Y∣X)P_t(Y\mid X) generally changes too; label shift is therefore not equivalent to covariate shift. (proceedings.mlr.press)

Concept shift describes changes in P(Y∣X)P(Y\mid X): the predictive meaning of the same observed features differs between settings. In sequential applications, concept drift often refers more broadly to distributional changes over time. Definitions differ, so “drift” does not always mean exclusively a changing conditional relationship. (rtg.cis.upenn.edu)

Causes and consequences

Shift may arise from nonrepresentative sampling, different measurement processes, geographical variation, or changing populations. Selection mechanisms can alter the distribution of observed examples even when the broader population remains unchanged. Changes between environments may also affect several distributional components simultaneously, rather than conforming to one idealized shift type. (sciencedirect.com)

In computer vision, differences between camera locations can change image backgrounds and viewing conditions. In natural language processing, deployment may involve different users or text domains. Remote sensing systems can encounter geographical and temporal differences. These are examples of application settings, not guarantees that a particular conditional distribution stays fixed. (proceedings.mlr.press)

Shift is particularly consequential when a model relies on correlations that do not persist across environments. Approaches informed by causal inference attempt to identify predictive relationships that remain stable under specified changes. Their validity depends on assumptions about the data-generating process and which mechanisms can change. (proceedings.mlr.press)

Detection and evaluation

Shift detection compares reference data with newly observed data. Methods include statistical hypothesis testing, tests after dimensionality reduction, and classifiers trained to distinguish source examples from target examples. Successful domain discrimination indicates distributional differences, but does not by itself measure the associated prediction error. (proceedings.neurips.cc)

Without target labels or additional assumptions, input monitoring cannot generally identify a change confined to P(Y∣X)P(Y\mid X). Nor does detecting changed inputs establish that a model has deteriorated: a shift can affect irrelevant features. Detection, characterization, and assessment of harm are therefore distinct tasks. (rtg.cis.upenn.edu)

Evaluation uses a test set designed to represent the intended target environment. Holding out locations, institutions, groups, or time periods can reveal weaknesses hidden by random splits. The WILDS benchmark, introduced in 2021, assembled ten datasets featuring naturally occurring shifts. Its experiments found substantial gaps between in-distribution and out-of-distribution performance, including for methods designed to address shift. (proceedings.mlr.press)

Adaptation and robustness

Under covariate shift and sufficient source coverage, importance weighting assigns examples weights proportional to

w(x)=pt(x)ps(x).w(x)=\frac{p_t(x)}{p_s(x)}.

This converts source expectations into target expectations. The approach requires suitable overlap: weighting cannot supply information about target regions absent from the source. Estimating large density ratios can also increase variance. (jmlr.csail.mit.edu)

Under label-shift assumptions, class-frequency ratios can instead support correction. Black Box Shift Estimation uses labeled source data, unlabeled target data, and a predictor’s confusion matrix to estimate changed class proportions, subject to identifiability conditions such as matrix invertibility. (proceedings.mlr.press)

Domain adaptation uses information from the target domain to improve transfer. One approach learns representations that support prediction while making source and target examples harder to distinguish. Distributionally robust optimization instead minimizes worst-case risk over a specified family of plausible distributions. Neither framework offers protection against unrestricted changes: adaptation depends on assumptions connecting domains, while robustness depends on which distributions its uncertainty set includes. (jmlr.csail.mit.edu)