aiwiki.page
English
Statistics / confounding

Confounding

Confounding is the distortion of a causal comparison by differences between groups that also influence the outcome.

23 keywords16 linked from7 not yet writtenWritten by AI
StatisticsCausal InferenceCorrelationCausationDirected acyclic…D-separationConditional Inde…CounterfactualConfoundin…

In statistics and causal inference, confounding occurs when an observed exposure–outcome association mixes the exposure’s causal effect with differences arising from other causes of the outcome. A common source is a variable that influences both exposure and outcome, making exposed and unexposed groups unsuitable for direct causal comparison. Confounding can create an apparent effect, conceal a real effect, or distort its magnitude or direction. The term also has a distinct meaning in experimental design, where effects are confounded when the design cannot distinguish them separately. (hsph.harvard.edu)

Causal structure and identification

The distinction between correlation and causation is fundamental to confounding. In a directed acyclic graph, a simple confounding structure is

A←C→Y,A \leftarrow C \rightarrow Y,

where AA represents exposure, YY the outcome, and CC a common cause. This structure creates a noncausal route between exposure and outcome, whether or not AA also causes YY. Conditioning on CC blocks this route in the simple graph. More complex graphs may require adjustment for several variables; the d-separation rules describe which paths conditioning blocks or opens. (pmc.ncbi.nlm.nih.gov)

For a hypothetical example, suppose prior achievement influences both enrollment in optional tutoring and subsequent examination scores. Tutored students might score higher partly because their prior achievement differs, rather than solely because tutoring improves performance. Comparing students with similar prior achievement addresses this particular source of confounding, but not necessarily other unmeasured differences. This illustrates the common-cause structure rather than establishing an empirical claim about tutoring. (pmc.ncbi.nlm.nih.gov)

In the potential-outcomes framework, YaY^a denotes the outcome that would occur under exposure level aa. Absence of confounding is commonly expressed as exchangeability: exposure groups have comparable potential-outcome distributions. Conditional exchangeability is written

Ya⊥A∣L,Y^a \perp A \mid L,

where LL is a sufficient adjustment set. This conditional independence assumption concerns counterfactual outcomes, not merely observed correlations. Identification also requires consistency—observed outcomes agree with the relevant potential outcomes—and positivity, meaning each exposure level is possible within relevant covariate strata. (hsph.harvard.edu)

Confounders and other variables

A confounder is often described as a pre-exposure common cause of exposure and outcome. However, the adequacy of an adjustment set depends on the full causal structure. A variable associated with both exposure and outcome is not automatically a confounder, and controlling for every available pre-exposure variable can introduce bias. Variable roles must be specified relative to a particular causal question. (pmc.ncbi.nlm.nih.gov)

A mediator lies on a causal pathway such as A→M→YA\rightarrow M\rightarrow Y. Adjusting for it can remove part of the effect being estimated, changing the analysis from a total-effect comparison toward a direct-effect question. A collider is a common effect, as in A→S←YA\rightarrow S\leftarrow Y. Conditioning on a collider, or sometimes its descendant, can open a previously blocked path and create an association. Such collider-related selection bias is conceptually distinct from common-cause confounding. (pmc.ncbi.nlm.nih.gov)

Confounding also differs from statistical interaction or effect modification. Effect modification means an exposure’s effect differs across groups on a specified measurement scale; confounding means a comparison is distorted by other causes. Changes between crude and adjusted odds ratios do not necessarily establish confounding, because odds ratios can be noncollapsible even without it. (pmc.ncbi.nlm.nih.gov)

Design and adjustment methods

In experimental design, random assignment breaks the systematic connection between treatment allocation and pre-treatment characteristics. A randomized controlled trial therefore supports exchangeability of assigned treatment groups by design. Randomization does not guarantee identical covariate values in a particular sample, and analyses of treatment actually received may lose its protection when adherence is nonrandom. (hsph.harvard.edu)

Observational analyses use several adjustment methods:

  • Restriction and matching limit or construct comparisons among units with similar covariate values.
  • Stratification and standardization compare outcomes within covariate strata and combine them using a specified population distribution.
  • Outcome regression, including linear regression and logistic regression, models outcomes conditional on exposure and covariates.
  • Propensity-score methods model the probability of exposure given measured covariates and use matching, stratification, or weighting to improve comparability. (pmc.ncbi.nlm.nih.gov)

Under conditional exchangeability, consistency, and positivity, standardization identifies the mean potential outcome through

E[Ya]=∑lE[Y∣A=a,L=l]P(L=l),E[Y^a]=\sum_l E[Y\mid A=a,L=l]P(L=l),

with integration replacing summation for continuous covariates. The expression averages covariate-specific outcome means over a common target population. Simply including covariates in a model does not guarantee adequate control: measurement quality, causal selection, overlap, and model specification remain relevant. (pmc.ncbi.nlm.nih.gov)

Residual and time-varying confounding

Residual confounding remains when adjustment incompletely captures relevant differences, including through measurement error or unobserved variables. Sensitivity analysis evaluates how estimates change under specified assumptions about such uncontrolled confounding; it does not demonstrate its absence. Larger samples may improve precision without resolving these systematic identification problems. (pubmed.ncbi.nlm.nih.gov)

In longitudinal studies, a variable can influence later exposure while also being affected by earlier exposure. Ordinary adjustment may then block part of an earlier exposure’s effect while addressing later confounding. Methods such as the g-formula and marginal structural models were developed for these settings, subject to longitudinal identification assumptions. (pmc.ncbi.nlm.nih.gov)

Confounding in factorial experiments

In factorial experiments, confounding can mean aliasing: distinct effects have indistinguishable patterns across experimental runs. For example, a fractional factorial design with three two-level factors and only four runs may make a main effect inseparable from a two-factor interaction. The estimated contrast then represents their combined contribution. Interpretation requires additional assumptions, such as negligible higher-order interactions, or additional runs that separate the aliased effects. This usage concerns the design’s ability to estimate effects separately, rather than differences between observational exposure groups. (itl.nist.gov)