A likelihood function is a function of a statistical model’s parameters obtained by holding the observed data fixed and evaluating their probability mass or probability density at different parameter values. It is fundamental to statistics, providing the basis for parameter estimation, model comparison, and uncertainty assessment. Unlike a probability distribution over possible observations, a likelihood function varies the parameters rather than the data. Its central interpretation is comparative: within the specified model, parameter values with higher likelihood assign greater probability mass or density to the observed data. (online.stat.psu.edu)
Definition and interpretation
Let (X) denote a random variable or data vector, and let (\theta\in\Theta) denote the model’s parameter, possibly a vector. If (f(x;\theta)) describes the sampling distribution, the likelihood after observing (X=x) is
[ L(\theta;x)=f(x;\theta). ]
For discrete data, (f) is a probability mass function; for continuous data, it is a probability density function. Thus, the same mathematical expression can describe either the sampling distribution or the likelihood, depending on which argument is allowed to vary. (online.stat.psu.edu)
A likelihood is not generally a probability distribution over (\theta): it need not sum or integrate to one over the parameter space. For continuous observations, it evaluates a density, not the probability of observing an exact point, which is ordinarily zero. Density-based likelihood values can exceed one. These distinctions prevent interpreting (L(\theta;x)) directly as the probability that (\theta) is correct. A probability distribution over parameters requires additional assumptions, as in Bayesian inference. (arxiv.org)
Constructing a sample likelihood
The likelihood of several observations is determined by their joint distribution. If (x_1,\ldots,x_n) arise from observations with statistical independence and a common sampling density or mass function, then
[ L(\theta;x_1,\ldots,x_n) =\prod_{i=1}^{n}f(x_i;\theta). ]
The product formula depends on the independence assumption; dependent observations require the appropriate joint model. Observations need not have identical distributions for independence to yield a product. (online.stat.psu.edu)
For example, suppose (n) independent trials follow a Bernoulli distribution with success probability (p). If the observed sequence contains (k) successes, its likelihood is
[ L(p)=p^k(1-p)^{n-k},\qquad 0\le p\le1. ]
If only the success count is recorded, its binomial distribution supplies the likelihood [ L_{\mathrm{count}}(p)=\binom nk p^k(1-p)^{n-k}. ] The binomial coefficient does not depend on (p), so both likelihoods have the same maximizer. More generally, multiplication by a positive factor independent of the parameter leaves parameter likelihood ratios and maximizers unchanged. (online.stat.psu.edu)
Log-likelihood and estimation
The log-likelihood is
[ \ell(\theta)=\log L(\theta). ]
Because the logarithm is strictly increasing, maximizing a positive likelihood is equivalent to maximizing its logarithm. Products become sums: [ \ell(\theta)=\sum_{i=1}^{n}\log f(x_i;\theta). ] This simplifies differentiation and numerical computation. Maximum likelihood estimation selects [ \hat\theta\in\operatorname*{arg,max}_{\theta\in\Theta}\ell(\theta). ] In the Bernoulli example, (\hat p=k/n) for (n>0), including boundary solutions when every outcome is identical. (itl.nist.gov)
Finding the maximum is a problem in mathematical optimization. Setting derivatives to zero can identify interior candidates, but does not by itself establish a global maximum. Constraints, boundary values, and multiple stationary points may matter. Some models have no finite maximizing parameter, and numerical procedures may fail to converge. These are properties of the model and optimization problem, not changes in the definition of likelihood. (itl.nist.gov)
Curvature and statistical inference
The gradient of the log-likelihood is called the score. Under suitable differentiability and regularity conditions, Fisher information can be expressed as the expected negative second derivative—or negative Hessian for a parameter vector—of the log-likelihood. It quantifies information about parameters supplied by the model’s observations. (arxiv.org)
Near a well-behaved maximum, log-likelihood curvature helps approximate an estimator’s variance and construct confidence intervals. Under regularity conditions and a correctly specified model, maximum likelihood estimators have an approximately normal distribution in large samples, with covariance related to inverse Fisher information. These approximations are not universal, particularly at parameter boundaries or in irregular models. (arxiv.org)
Likelihood also supports statistical hypothesis testing. For a restricted parameter space (\Theta_0\subseteq\Theta), a likelihood-ratio statistic compares [ \Lambda= \frac{\sup_{\theta\in\Theta_0}L(\theta)} {\sup_{\theta\in\Theta}L(\theta)}. ] Small values indicate poorer fit under the restriction. Under appropriate regularity conditions, (-2\log\Lambda) has an asymptotic chi-square distribution, with degrees of freedom corresponding to the number of independent restrictions. (itl.nist.gov)
Bayesian inference and machine learning
In Bayesian inference, Bayes’ theorem combines the likelihood with a prior distribution (\pi(\theta)) to obtain a posterior distribution: [ \pi(\theta\mid x)= \frac{L(\theta;x)\pi(\theta)} {\int_\Theta L(u;x)\pi(u),du}, ] when the denominator is finite and positive. This denominator is the marginal likelihood. The posterior incorporates both the sampling model and the prior; the likelihood alone does not. (online.stat.psu.edu)
In machine learning, negative log-likelihood frequently serves as a loss function. For categorical prediction, it yields the familiar cross-entropy objective. In linear regression with independent Gaussian errors of constant variance, maximizing likelihood over regression coefficients is equivalent to minimizing squared error. These equivalences depend on the assumed probabilistic model. (web.stanford.edu)