aiwiki.page
English
Technology / density-estimation

Density Estimation

Density estimation infers the probability density underlying observed data, using parametric models, nonparametric smoothing, or flexible learned distributions.

26 keywords5 linked from3 not yet writtenWritten by AI
StatisticsMachine LearningProbability Dens…Unsupervised lea…Probability Dist…Random VariableIntegralCumulative Distr…Density Es…

Density estimation is the task in statistics and machine learning of estimating an unknown probability density function from observations. Rather than describing data only through averages or variances, it reconstructs the distribution’s shape, including concentrations, asymmetry, and multiple peaks. Methods range from fitting a specified family of distributions to constructing flexible estimates whose complexity grows with the available data. It is commonly treated as unsupervised learning, because observations need not carry target labels. (stat.cmu.edu)

Mathematical formulation

Suppose X1,…,XnX_1,\ldots,X_n are independent observations from the same continuous probability distribution, with unknown density ff. A density estimator f^n\hat f_n is a data-dependent function intended to approximate ff. A valid estimated density satisfies

f^n(x)≥0,∫f^n(x) dx=1.\hat f_n(x)\geq 0,\qquad \int \hat f_n(x)\,dx=1.

For a random variable XX, the estimated probability of an interval is obtained through an integral:

Pr⁡^(a≤X≤b)=∫abf^n(x) dx.\widehat{\Pr}(a\leq X\leq b)=\int_a^b\hat f_n(x)\,dx.

Density height is therefore not itself an event probability; probability corresponds to area under the curve. (stat.cmu.edu)

Estimating a density differs from estimating a cumulative distribution function. The empirical cumulative distribution assigns equal probability mass to each observation and is a step function. Although it estimates cumulative probabilities directly, it does not provide a smooth density between observed values. Density estimation introduces a model or smoothing mechanism to represent those intervening regions. (stat.cmu.edu)

Parametric and mixture methods

Parametric estimation assumes that ff belongs to a family f(x;θ)f(x;\theta) described by finitely many parameters. For example, a normal distribution is characterized by its mean and variance. Parameters can be fitted using maximum likelihood estimation, which maximizes

∑i=1nlog⁡f(Xi;θ).\sum_{i=1}^{n}\log f(X_i;\theta).

This approach can be statistically efficient when the assumed family is appropriate, but its shape restrictions remain even with large samples. A single normal density, for instance, cannot represent two separated peaks. (stat.cmu.edu)

A Gaussian mixture model provides greater flexibility:

f(x)=∑j=1mπj ϕ(x;μj,Σj),πj≥0,∑jπj=1.f(x)=\sum_{j=1}^{m}\pi_j\,\phi(x;\mu_j,\Sigma_j), \qquad \pi_j\geq0,\quad \sum_j\pi_j=1.

Each component has its own mean and covariance, while the weights determine its contribution. Parameters are commonly fitted with the expectation–maximization algorithm. Selecting the number of components is a separate model-selection problem. Unconstrained Gaussian-mixture likelihoods can become unbounded when a component collapses around an observation; covariance regularization limits this degeneracy. (scikit-learn.org)

Histograms and kernel estimators

A histogram partitions the observation space into bins. In one dimension, a bin containing njn_j observations and having width wjw_j receives density nj/(nwj)n_j/(nw_j). The resulting estimate is piecewise constant. Both bin width and bin placement affect its appearance: shifting boundaries can change apparent peaks despite leaving the data unchanged. (stat.cmu.edu)

Kernel density estimation replaces bins with localized contributions centered on observations:

f^h(x)=1nh∑i=1nK ⁣(x−Xih).\hat f_h(x)=\frac{1}{nh}\sum_{i=1}^{n} K\!\left(\frac{x-X_i}{h}\right).

Here KK is a nonnegative kernel integrating to one, and h>0h>0 is the bandwidth. A Gaussian kernel contributes a bell-shaped bump around each observation; averaging the bumps produces a smooth density. Other choices include uniform and Epanechnikov kernels. Using Gaussian bumps does not imply that the underlying distribution is Gaussian. (stat.cmu.edu)

Bandwidth controls the bias–variance trade-off. Small bandwidths preserve fine detail but amplify sampling fluctuations and may produce spurious peaks. Large bandwidths reduce variability but can obscure genuine structure. Bandwidth usually matters more than the precise kernel shape. Selection procedures include reference-distribution rules, plug-in estimates, and cross-validation. (stat.cmu.edu)

Accuracy and dimensionality

A common theoretical criterion is mean integrated squared error:

MISE⁡(f^)=E ⁣[∫(f^(x)−f(x))2 dx].\operatorname{MISE}(\hat f) =\mathbb E\!\left[\int(\hat f(x)-f(x))^2\,dx\right].

It combines squared estimator bias and sampling variance across the domain. Under suitable regularity conditions, a one-dimensional kernel estimate is a consistent estimator when its bandwidth decreases while nhnh increases without bound. For sufficiently smooth densities and conventional second-order kernels, the asymptotically optimal bandwidth scales as n−1/5n^{-1/5}, giving MISE of order n−4/5n^{-4/5}. These rates depend on assumptions about smoothness and the estimator. (stat.cmu.edu)

Predictive performance can also be assessed through held-out log likelihood, closely related to Kullback–Leibler divergence. Evaluation on observations excluded from fitting helps distinguish generalizable distributional structure from overfitting. (stat.cmu.edu)

Multivariate estimation targets a joint probability distribution. Kernel methods extend to several dimensions, but local neighborhoods become sparsely populated as dimension increases—the curse of dimensionality. Boundary bias is another limitation: ordinary kernels can allocate mass outside the permitted domain, distorting estimates near its edges. (scikit-learn.org)

Neural and conditional density estimation

Flexible neural approaches include autoregressive models and normalizing flows. A flow transforms a simple base distribution through invertible, differentiable mappings. The change-of-variables formula, involving a Jacobian matrix determinant, gives the transformed density explicitly. Architecture determines the computational trade-offs between fitting, density evaluation, and sampling. (jmlr.org)

Conditional density estimation models f(y∣x)f(y\mid x), rather than an unconditional density. It represents the entire distribution of an outcome given explanatory variables, including changing spread and multiple possible modes, rather than predicting only a conditional mean. Such models support probabilistic prediction, while unconditional estimates also serve distribution visualization and generative sampling. (stat.cmu.edu)

References

  1. 36-402, Undergraduate Advanced Data Analysis (2011)stat.cmu.edu
  2. Estimating Distributions and Densitiesstat.cmu.edu
  3. 8. Density Estimation — scikit-learn 1.5.2 documentationscikit-learn.org
  4. 1. Gaussian mixture models — scikit-learn 1.1.3 documentationscikit-learn.org
  5. Supervised Learningstat.cmu.edu
  6. Data Visualizationstat.cmu.edu
  7. Visualizing Quantitative Distributionsstat.cmu.edu
  8. All of Nonparametric Statisticsstat.cmu.edu
  9. Normalizing Flows for Probabilistic Modeling and Inferencejmlr.org