Bayesian inference is an approach to statistics that uses Bayes’ theorem to update uncertainty about unknown quantities after observing data. It represents uncertainty through probability distributions, combines an initial distribution with a model of the observations, and obtains an updated distribution. Its central output is therefore not necessarily a single estimate, but a distribution supporting estimation, prediction, and comparisons between hypotheses. All such conclusions are conditional on the specified model and prior information. (sites.stat.columbia.edu)
Mathematical framework
Let (D) denote observed data and (\theta) an unknown parameter, possibly a vector. Bayesian inference expresses the relationship between these quantities as
[ p(\theta\mid D)= \frac{p(D\mid\theta)p(\theta)}{p(D)}. ]
The prior distribution (p(\theta)) describes uncertainty before incorporating (D). The likelihood function (p(D\mid\theta)) describes how the statistical model assigns probabilities or densities to the observations, considered as a function of (\theta). The posterior distribution (p(\theta\mid D)) describes uncertainty after conditioning on the data. Computing the posterior amounts to multiplying prior and likelihood and normalizing their product. (bayesball.github.io)
For continuous parameters, the normalizing quantity is the marginal likelihood, or model evidence:
[ p(D)=\int p(D\mid\theta)p(\theta),d\theta. ]
For discrete parameters, a sum replaces the integral. Provided the normalizing quantity is finite and positive, the posterior integrates or sums to one. The shorthand (p(\theta\mid D)\propto p(D\mid\theta)p(\theta)) omits this parameter-independent denominator. (sites.stat.columbia.edu)
Unlike a posterior distribution, a likelihood need not integrate to one over parameter values. Confusing (p(D\mid\theta)) with (p(\theta\mid D)) reverses the conditioning: a model that makes observations probable does not, by itself, specify the probability of that model after those observations. Prior weights also matter. (bayesball.github.io)
A binary-outcome example
Suppose independent trials have an unknown success probability (\theta). Each outcome follows a Bernoulli distribution, and (s) successes among (n) trials give a likelihood proportional to
[ \theta^s(1-\theta)^{n-s}. ]
Choose a beta distribution as the prior, with positive parameters (\alpha) and (\beta). Multiplying its density by the likelihood yields
[ \theta\mid D\sim \operatorname{Beta}(\alpha+s,\beta+n-s). ]
The beta family is thus a conjugate prior family for this likelihood: updating preserves the distributional family while changing its parameters. (online.stat.psu.edu)
As a worked illustration, a uniform prior (\operatorname{Beta}(1,1)) and seven successes in ten trials produce (\operatorname{Beta}(8,4)). Its posterior mean is (8/12=2/3), whereas maximum likelihood estimation gives (7/10). Under this model, the probability of success on the next trial equals the posterior mean. These results follow directly from the beta update and predictive averaging; they do not establish that the trials actually satisfy the independence and constant-probability assumptions. (online.stat.psu.edu)
Estimation and prediction
Posterior distributions can be reported through means, medians, quantiles, or a maximum a posteriori estimate, which selects a posterior mode. A point estimate discards information about spread, asymmetry, dependence, and multiple modes, so it is not equivalent to the complete posterior. (mc-stan.org)
A credible interval contains a specified posterior probability. A 95% credible interval assigns probability 0.95 to the parameter lying within its endpoints, conditional on the model, prior, and observed data. This differs from a confidence interval, whose frequentist interpretation concerns the coverage of an interval-producing procedure over repeated samples. Numerical agreement between intervals does not make their interpretations identical. (sites.stat.columbia.edu)
Prediction incorporates uncertainty about parameters rather than treating an estimate as known. Assuming a future observation (\tilde y) is conditionally independent of the observed data given (\theta), the posterior predictive distribution is
[ p(\tilde y\mid D)= \int p(\tilde y\mid\theta)p(\theta\mid D),d\theta. ]
This combines uncertainty about the parameter with randomness in future observations. Simulation typically draws a parameter from the posterior and then draws a future observation from the corresponding data model. (mc-stan.org)
Computation
Conjugate models sometimes permit exact analytic updates, but more complex models require numerical computation. Markov chain Monte Carlo constructs dependent draws whose distribution approaches the posterior under suitable conditions. Posterior means and other quantities can then be estimated from these draws. Convergence diagnostics, effective sample sizes, and Monte Carlo error estimates assess computational reliability; a large number of draws alone does not establish adequate exploration. (mc-stan.org)
Variational inference instead selects an approximating distribution from a tractable family using mathematical optimization. Common formulations minimize Kullback–Leibler divergence, equivalently maximizing an evidence lower bound. This can reduce computational cost, but restricts results to what the approximation family can represent. Laplace approximation provides another approach, using a local normal distribution approximation around a mode. (arxiv.org)
Model checking and limitations
A mathematically valid posterior does not demonstrate that a model adequately describes its application. Posterior predictive checking compares observed features with replicated data generated from the fitted model. Discrepancies in tails, variability, trends, or other relevant features can reveal deficiencies. Such checks examine particular aspects of fit rather than proving a model correct. (sites.stat.columbia.edu)
Sensitivity analysis examines how conclusions change under plausible alternative priors and modeling assumptions. Cross-validation evaluates predictive performance on observations excluded from fitting. Model checking and computational checking address different problems: an accurately computed posterior may belong to an inadequate model, while a useful model may still be poorly explored by its inference algorithm. (sites.stat.columbia.edu)