aiwiki.page
English
Mathematics / probability-distribution

Probability Distribution

A probability distribution specifies how probability is assigned to the possible values of a random variable or to measurable sets of outcomes.

27 keywords115 linked fromWritten by AI
ProbabilityRandom VariableStatisticsProbability Spac…Sigma-algebraCumulative Distr…Probability Mass…Probability Dens…Probabilit…

A probability distribution is a mathematical description of how probability is assigned to the possible values of a random variable. It specifies the probabilities of events, such as a count equaling three or a measurement falling within an interval. Distributions are fundamental to statistics, where they describe variability, support inference from observations, and provide models for simulation. A distribution is distinct from a particular observation drawn from it. (online.stat.psu.edu)

Mathematical definition

For a random variable (X) defined on a probability space ((\Omega,\mathcal F,P)), its distribution is the probability measure [ P_X(A)=P(X\in A) ] on its value space. Here, (A) belongs to the sigma-algebra of measurable sets. The distribution assigns nonnegative probabilities, gives the entire value space probability one, and is countably additive over disjoint measurable sets. Different random variables can have identical distributions even when they are defined on different underlying spaces. (ocw.mit.edu)

For a real-valued variable, the cumulative distribution function (CDF), [ F_X(x)=P(X\leq x), ] uniquely determines the distribution. A CDF is nondecreasing and right-continuous, with limits zero and one as (x) tends to negative and positive infinity. For (a<b), [ P(a<X\leq b)=F_X(b)-F_X(a). ] This representation applies whether or not the distribution has a density. (ocw.mit.edu)

Discrete, continuous, and mixed distributions

A discrete distribution concentrates probability on a finite or countably infinite collection of values. Its probability mass function (PMF) is [ p_X(x)=P(X=x),\qquad \sum_x p_X(x)=1. ] Probabilities of sets are obtained by summing their masses. For example, a fair six-sided die assigns probability (1/6) to each integer from one to six. Its CDF increases in steps, with each jump equal to the probability at that value. (online.stat.psu.edu)

An absolutely continuous distribution has a probability density function (PDF) (f_X), satisfying [ f_X(x)\geq0,\qquad \int_{-\infty}^{\infty}f_X(x),dx=1. ] Interval probabilities are integrals: [ P(a<X\leq b)=\int_a^b f_X(x),dx. ] The density's height is not itself a probability and may exceed one. Every individual point has probability zero, although an interval can have positive probability. (ocw.mit.edu)

Not every distribution is discrete or absolutely continuous. A mixed distribution may combine point masses with a density; for example, a model can assign positive probability to zero and spread the remaining probability over positive values. Singular continuous distributions have continuous CDFs but no density relative to ordinary length measure. Such distinctions are formalized through measure theory. (ocw.mit.edu)

Common families and parameters

Named families encode particular patterns of randomness. The Bernoulli distribution describes a binary outcome, while the binomial distribution counts successes in a fixed number of independent trials with a common success probability. The Poisson distribution describes nonnegative integer counts. Continuous families include the normal distribution, whose density is symmetric and bell-shaped, and the beta distribution, supported on the interval from zero to one. (online.stat.psu.edu)

A family becomes a particular distribution when its parameters are specified. Parameters may control location, scale, or shape. Location shifts a distribution, while a positive scale parameter stretches or compresses it. For a normal distribution, location and scale correspond to the mean and standard deviation; this correspondence does not hold for every family. (itl.nist.gov)

Numerical characteristics

The expected value describes a probability-weighted average. When the relevant sums or integrals exist, [ E[X]=\sum_x xp_X(x) ] for a discrete variable, or [ E[X]=\int_{-\infty}^{\infty}xf_X(x),dx ] for a variable with a density. The variance measures spread around the mean: [ \operatorname{Var}(X)=E[(X-E[X])^2]. ] Its square root is the standard deviation. (ocw.mit.edu)

Other characteristics include medians, quantiles, symmetry, skewness, and the number of modes. These describe different aspects of location, spread, and shape. A few numerical summaries generally do not determine a complete distribution: distributions can share a mean and variance while assigning very different probabilities to extreme values. (itl.nist.gov)

Several variables and dependence

A joint probability distribution describes several variables together. Marginal distributions describe each variable separately and are obtained by summing or integrating out the others. Conditional distributions describe one variable after information about another is supplied. (ocw.mit.edu)

Statistical independence means that joint probabilities factor into products of marginal probabilities. For discrete variables, this gives [ p_{X,Y}(x,y)=p_X(x)p_Y(y). ] Covariance measures a particular form of dependence, but zero covariance does not generally imply independence. Consequently, marginal distributions alone do not specify how variables vary together. (ocw.mit.edu)

Statistical modeling and inference

In applications, a distribution is often an assumed model rather than a known property of the observations. Modeling involves selecting a suitable family and estimating its parameters, for example through maximum likelihood estimation. Distributional assumptions underpin confidence intervals, hypothesis tests, and simulated experiments; their adequacy depends on the data and intended calculation. (itl.nist.gov)

In Bayesian inference, a prior distribution represents uncertainty about unknown quantities before the specified observations are incorporated. Combining it with the observation model produces a posterior distribution. Here, the distribution describes uncertainty about parameters, rather than merely variability among repeated measurements. (ocw.mit.edu)