Maximum a posteriori estimation, usually abbreviated MAP estimation, is a method of Bayesian inference that represents an unknown parameter by a value maximizing its posterior distribution. It combines evidence supplied by observations with a prior distribution over possible parameter values. Unlike methods that retain the entire posterior, MAP produces a point estimate. It is closely related to maximum likelihood estimation, but incorporates the prior into the optimization objective. (cs.cmu.edu)
Mathematical definition
Let denote observed data and a parameter belonging to a parameter space . By Bayes’ theorem,
Here , considered as a function of , is the likelihood function, while is the marginal likelihood. Provided the posterior is well defined, a MAP estimate satisfies
The membership notation allows for several equally maximizing values. Because the denominator is independent of , it can be omitted during maximization. Wherever the relevant densities are positive, taking logarithms gives the equivalent objective function
For discrete parameters, MAP maximizes posterior probability mass. For continuous parameters, it maximizes a probability density function, not the probability of an exact value: individual points ordinarily have zero probability. The estimate therefore identifies a density peak, rather than a region containing most of the posterior probability. (mc-stan.org)
Relationship to maximum likelihood and regularization
Maximum likelihood selects parameters using only . MAP selects them using the product . A prior that is constant throughout the relevant parameter space leaves the maximizing values unchanged. However, a constant density on an unbounded space is generally not a proper probability distribution, so the equivalence requires care about the prior’s support and normalization. (cs229.stanford.edu)
Minimizing the negative log posterior makes the connection with regularization explicit:
The first term measures disagreement with the observations; the second acts as a parameter penalty. An independent, zero-mean Gaussian prior produces a squared penalty. Independent, zero-centered Laplace distributions produce an penalty. Thus familiar regularized objectives in machine learning can have a MAP interpretation under specified probabilistic assumptions. (cs229.stanford.edu)
For example, in linear regression with independent Gaussian noise of known variance and a Gaussian coefficient prior with covariance ,
This is ridge regression, with penalty strength determined by the ratio of noise variance to prior variance. If the data-fitting term is averaged rather than summed, the numerical penalty coefficient also depends on sample size. A Laplace coefficient prior instead yields lasso regression. (www2.stat.duke.edu)
Example: estimating a Bernoulli probability
Suppose independent observations follow a Bernoulli distribution with success probability , and successes are observed. Assign a beta prior,
This is a conjugate prior: the posterior belongs to the same distribution family,
When both posterior shape parameters exceed one, its unique interior mode is
This formula follows by differentiating the log posterior; outside these conditions, boundary behavior or nonuniqueness must be considered instead. (cs.cmu.edu)
For a prior and eight successes in ten trials, the formula gives , compared with the maximum likelihood estimate . The posterior mean is instead
which gives , approximately . These values illustrate that a posterior’s mode and expected value need not coincide. (cs.cmu.edu)
Computation and interpretation
MAP estimation is a mathematical optimization problem. Simple conjugate models may admit analytic solutions; more complicated models require numerical methods. Smooth objectives can be optimized using Newton’s method or quasi-Newton algorithms such as BFGS and L-BFGS. Numerical termination does not itself establish that a global maximum has been found, especially for multimodal posteriors. (mc-stan.org)
Within decision theory, the preferred point estimate depends on the loss function. For a discrete parameter and zero–one loss, choosing a posterior mode minimizes posterior expected loss. Under squared-error loss, the minimizing estimate is the posterior mean, not generally MAP. For continuous parameters, exact zero–one loss does not distinguish candidate values; interpreting MAP through shrinking neighborhoods requires additional regularity conditions. (hsong1.github.io)
Limitations
Continuous MAP estimates depend on parameterization. For an invertible differentiable transformation , the transformed posterior density includes a factor involving the Jacobian:
Consequently, transforming a MAP estimate need not yield the MAP estimate in the new coordinates, even though the underlying posterior probability measure is unchanged. (mc-stan.org)
A single mode also omits posterior spread, asymmetry, dependence, and alternative modes. Substituting MAP parameters into a predictive model is generally different from computing the posterior predictive distribution, which averages predictions over parameter uncertainty. MAP therefore supplies a particular point summary, not a complete description of Bayesian uncertainty. (cs229.stanford.edu)