Posterior predictive checking is a method of model assessment in Bayesian inference. It compares observed data with hypothetical replicated data generated under a fitted model, asking whether the model reproduces features relevant to the analysis. Comparisons may be graphical or numerical; their purpose is to identify particular discrepancies between model and data, rather than establish that a model is true. (sites.stat.columbia.edu)
Mathematical formulation
Let denote observed data and the model parameters. A Bayesian model specifies a prior distribution and a sampling model , which also determines the likelihood function. Conditioning on gives the posterior distribution . Under conditional replication, the posterior predictive distribution is
Here is a hypothetical repetition of the data-generating process. Its definition includes what is held fixed, such as sample size or experimental design. (sites.stat.columbia.edu)
This distribution incorporates both uncertainty about the parameters and variation in observations conditional on those parameters. Simulating observations only at a single parameter estimate generally omits the first component. (mc-stan.org)
Simulation and comparison
A simulation-based check draws
then compares observed and replicated features. A test statistic may depend only on the data; a more general discrepancy may also depend on unknown parameters. For a parameter-dependent discrepancy, each replicated value is compared with the observed discrepancy evaluated at the same posterior draw. (sites.stat.columbia.edu)
Graphical comparisons can overlay observed and replicated distributions or display the observed statistic against its simulated reference distribution. Possible features include the sample mean, standard deviation, quantiles, and extreme values. The choice determines which aspects of fit the check can reveal; agreement on one feature does not imply agreement on others. (mc-stan.org)
For example, a Poisson distribution constrains its mean and variance to be equal. A fitted Poisson model can reproduce the mean of overdispersed count data while generating substantially less variation than observed. Checking dispersion therefore reveals a failure that checking the mean alone may miss. (mc-stan.org)
Posterior predictive p-values
A numerical summary is the posterior predictive p-value, commonly defined as
Its Monte Carlo estimate is
This is a conditional probability of the specified comparison, not the probability that the model is correct. (arxiv.org)
Values near either endpoint can identify an unusually positioned observed discrepancy. However, unlike a classically calibrated continuous-null p-value, a posterior predictive p-value generally does not have a uniform distribution under repeated sampling from a correctly specified model. Ordinary hypothesis-testing thresholds therefore do not automatically retain their usual error-rate interpretation. (mc-stan.org)
Interpretation and limitations
The observations serve twice: first to fit the posterior, then as the reference for checking predictions. This reuse can make checks insensitive to some departures, especially features already accommodated during fitting. Posterior predictive p-values often concentrate toward the middle of their range, but describing them as universally conservative is mathematically incorrect. (arxiv.org)
A satisfactory check establishes only that the selected comparison has not exposed a discrepancy. It is not, by itself, an assessment of prediction on unseen data. Held-out predictive checks and cross-validation instead evaluate observations excluded from the corresponding fit, addressing a different assessment question. (arxiv.org)
Related predictive checks
Prior predictive checking generates parameters from the prior and then generates observations from the sampling model, without conditioning on observed outcomes. It examines the observable implications of the model and prior before fitting. (mc-stan.org)
For hierarchical models, replication may retain fitted group-specific parameters or generate new group parameters from the population model. These alternatives check different levels of the model: behavior within existing groups versus variation across newly generated groups. Mixed predictive checks combine posterior draws of population-level parameters with new draws of group-level parameters. (arxiv.org)
References
- Posterior Predictive Assessment of Model Fitness via Realized Discrepanciessites.stat.columbia.edu
- Posterior and Prior Predictive Checksmc-stan.org
- Posterior Predictive Samplingmc-stan.org
- Posterior predictive p-values and the convex orderarxiv.org
- Population Predictive Checksarxiv.org