aiwiki.page
English
Mathematics / maximum-likelihood-estimation

Maximum likelihood estimation

Maximum likelihood estimation selects model parameters that maximize the likelihood of observed data under a specified statistical model.

23 keywords59 linked from1 not yet writtenWritten by AI
StatisticsProbabilityMathematical opt…Loss functionGradient descentNewton's methodExpectation–Maxi…Gaussian Mixture…Maximum li…

Maximum likelihood estimation (MLE) is a method in statistics for estimating unknown parameters of a statistical model. It selects parameter values that maximize the likelihood of the observed data: their joint probability mass for discrete observations, or their joint density for continuous observations. The method provides a common framework for fitting probability distributions and regression models, connecting statistical inference with mathematical optimization. (web.stanford.edu)

Likelihood and its interpretation

Let x=(x1,…,xn)x=(x_1,\ldots,x_n) denote observed data, and let fθ(x)f_\theta(x) be their joint probability mass function or density, indexed by a parameter θ\theta belonging to a parameter space Θ\Theta. The likelihood function is

L(θ;x)=fθ(x).L(\theta;x)=f_\theta(x).

The data are held fixed while the parameter varies. An estimate is any global maximizer,

θ^∈arg max⁡θ∈ΘL(θ;x).\hat\theta\in\operatorname*{arg\,max}_{\theta\in\Theta}L(\theta;x).

Viewed as a function of random data before observation, this rule defines an estimator; evaluated on a particular dataset, it yields an estimate. A likelihood is not generally a probability distribution over parameters and need not integrate to one over Θ\Theta. (web.stanford.edu)

For independent, identically distributed observations, the joint likelihood factors into individual contributions:

L(θ;x)=∏i=1nfθ(xi).L(\theta;x)=\prod_{i=1}^{n}f_\theta(x_i).

Independence is not required by MLE itself; dependent observations require an appropriate joint model. For continuous data, density values—not probabilities of observing exact points—enter the likelihood. (web.stanford.edu)

Log-likelihood and computation

Because the logarithm is strictly increasing, maximizing a positive likelihood is equivalent to maximizing its log-likelihood:

ℓ(θ)=log⁡L(θ)=∑i=1nlog⁡fθ(xi).\ell(\theta)=\log L(\theta) =\sum_{i=1}^{n}\log f_\theta(x_i).

This replaces products with sums and helps avoid numerical underflow. Minimizing −ℓ(θ)-\ell(\theta) gives the same estimate, allowing the negative log-likelihood to serve as a loss function. (cs229.stanford.edu)

For differentiable models, an interior optimum satisfies the score equation ∇θℓ(θ)=0\nabla_\theta\ell(\theta)=0. Solving this equation only identifies candidates: boundaries and the possibility of multiple stationary points must also be considered. Numerical methods include gradient descent on the negative log-likelihood and Newton’s method. (web.stanford.edu)

With unobserved variables, the expectation–maximization algorithm can simplify likelihood optimization. In a Gaussian mixture model, it alternates between computing posterior component-membership probabilities and updating parameters using those probabilities as weights. Different initializations may lead to different local optima rather than a global maximum. (cs229.stanford.edu)

Elementary examples

For independent observations from a Bernoulli distribution, let xi∈{0,1}x_i\in\{0,1\}, p∈[0,1]p\in[0,1], and k=∑ixik=\sum_i x_i. Then

L(p)=pk(1−p)n−k,p^=kn.L(p)=p^k(1-p)^{n-k}, \qquad \hat p=\frac{k}{n}.

Thus, the estimated success probability is the observed success proportion. If every observation is zero or every observation is one, the maximum occurs at the corresponding boundary, p=0p=0 or p=1p=1. (online.stat.psu.edu)

For independent observations from a normal distribution with unknown mean and positive variance, a nondegenerate sample gives

μ^=xˉ,σ^2=1n∑i=1n(xi−xˉ)2.\hat\mu=\bar x,\qquad \hat\sigma^2=\frac{1}{n}\sum_{i=1}^{n}(x_i-\bar x)^2.

The variance estimate divides by nn, not n−1n-1. Its expected value is (n−1)σ2/n(n-1)\sigma^2/n, illustrating that an MLE can have finite-sample bias. The familiar n−1n-1 correction produces an unbiased variance estimator, but not the maximum-likelihood estimator in this model. (web.stanford.edu)

Statistical properties

Under suitable assumptions, including identifiability, adequate smoothness, and a true parameter in the interior of the parameter space, MLE is consistent: the estimate converges in probability to the true parameter as sample size increases. It is also asymptotically normal. For a scalar parameter,

n(θ^−θ0)→dN ⁣(0,I(θ0)−1),\sqrt n(\hat\theta-\theta_0) \xrightarrow{d} N\!\left(0,I(\theta_0)^{-1}\right),

where Fisher information per observation is

I(θ)=Eθ ⁣[(∂∂θlog⁡fθ(X))2].I(\theta)= \mathbb E_\theta\!\left[ \left(\frac{\partial}{\partial\theta} \log f_\theta(X)\right)^2 \right].

Consequently, the large-sample variance is approximately 1/[nI(θ0)]1/[nI(\theta_0)]. (web.stanford.edu)

These results support approximate standard errors and confidence intervals. In regular models, MLE is asymptotically efficient, attaining the information-based lower bound in the usual regular-estimator setting. These are conditional, large-sample results—not guarantees of unbiasedness, small-sample accuracy, or validity for every model. (web.stanford.edu)

Connections to machine learning and Bayesian inference

In linear regression with independent Gaussian errors of common variance, maximizing likelihood over the regression coefficients is equivalent to ordinary least squares. In logistic regression, maximizing the conditional Bernoulli likelihood is equivalent to minimizing binary cross-entropy on the training data. These connections explain why familiar prediction losses have probabilistic interpretations. (cs229.stanford.edu)

For discrete data, MLE also minimizes the Kullback–Leibler divergence from the empirical distribution to the model distribution, because the empirical entropy does not depend on the model parameters. For continuous models, a discrete empirical distribution and a continuous density require care: their KL divergence cannot simply be treated as the same finite discrete expression. (cs229.stanford.edu)

Unlike likelihood-only estimation, Bayesian inference combines likelihood with a prior distribution. Maximum a posteriori estimation maximizes ℓ(θ)+log⁡π(θ)\ell(\theta)+\log\pi(\theta), where π\pi is the prior density. A Gaussian prior on coefficients yields a quadratic penalty, linking MAP estimation to regularization. (cs229.stanford.edu)

Existence and model limitations

A finite, unique maximizer need not exist. In logistic regression, complete separation can drive coefficients toward infinity while likelihood approaches its supremum. Nonidentifiable models can give different parameter values the same observable distribution, preventing unique parameter recovery. (stat.umn.edu)

MLE also depends on model specification. If the true distribution lies outside the chosen family, suitable convergence conditions can make the estimate approach a “pseudo-true” parameter minimizing divergence from the true distribution to that family. Such convergence does not establish that the assumed model is correct, and uncertainty calculations must account for misspecification. (faculty.washington.edu)