Expected value is the probability-weighted average of a random variable. Usually written (\mathbb E[X]), (E[X]), or (\mu), it describes the center of a probability distribution by weighting possible values according to their probabilities. It is a theoretical quantity determined by a probability model, not necessarily an observable outcome or a prediction of the most likely result. Expectation is a foundational concept in statistics and probability theory. (statlect.com)
Definition and calculation
For a discrete random variable (X) with possible values (x_i) and probabilities (p_i=P(X=x_i)), its finite expected value is
[ \mathbb E[X]=\sum_i x_i p_i, ]
provided that (\sum_i |x_i|p_i<\infty). For finitely many finite-valued outcomes, this condition holds automatically. The weights sum to one, so expectation is a weighted arithmetic average. (statlect.com)
For an absolutely continuous random variable with probability density (f_X(x)),
[ \mathbb E[X]=\int_{-\infty}^{\infty}x f_X(x),dx, ]
provided that the corresponding absolute-value integral is finite. A density is not the probability of an individual value: probabilities are obtained by integrating the density over sets. Thus the continuous formula weights values by probability density rather than by point probabilities. (statlect.com)
In measure theory, both formulas are instances of
[ \mathbb E[X]=\int_\Omega X(\omega),dP(\omega), ]
where (\Omega) is the sample space and (P) is its probability measure. This definition also covers distributions that are neither discrete nor absolutely continuous. A finite expectation exists exactly when (\mathbb E[|X|]<\infty). Nonnegative variables may have expectation (+\infty). For a signed variable, if its positive and negative parts both have infinite expectations, their difference is undefined—not zero. (ocw.mit.edu)
Interpretation and examples
For a fair six-sided die, the formula gives
[ \mathbb E[X]=\frac{1+2+3+4+5+6}{6}=3.5. ]
No individual roll produces 3.5. This illustrates why an expected value need not be a possible outcome. For a Bernoulli distribution, which assigns probability (p) to 1 and (1-p) to 0, the expectation is (p). In particular, the expectation of an indicator variable recording whether an event occurred equals that event’s probability. (ocw.mit.edu)
The long-run interpretation is formalized by the law of large numbers. If (X_1,X_2,\ldots) are independent, identically distributed random variables with finite absolute expectation, their sample average
[ \overline X_n=\frac1n\sum_{i=1}^{n}X_i ]
converges almost surely to their common expectation. This is a statement about averages as the number of observations increases; it does not guarantee a particular result in any finite experiment. (ocw.mit.edu)
For independent observations with common finite variance (\sigma^2), the sample average has variance (\sigma^2/n). Under these assumptions, the central limit theorem further describes the asymptotic distribution of appropriately scaled fluctuations around the expectation. It complements, rather than defines, the meaning of expectation. (ocw.mit.edu)
Algebraic properties
The central computational property is linearity. For integrable (X) and (Y) and constants (a,b,c),
[ \mathbb E[aX+bY+c] =a\mathbb E[X]+b\mathbb E[Y]+c. ]
Independence is unnecessary. Consequently, the expected total of several quantities can be calculated by adding their expectations even when the quantities are dependent. By contrast, (\mathbb E[XY]=\mathbb E[X]\mathbb E[Y]) holds for independent integrable variables but not generally for dependent ones. (ocw.mit.edu)
Expectation also preserves order: if (X\le Y) almost surely, then (\mathbb E[X]\le\mathbb E[Y]), whenever the expectations are finite. However, a nonlinear function generally cannot be moved outside the expectation. For a convex function (g), Jensen’s inequality gives
[ g(\mathbb E[X])\le \mathbb E[g(X)] ]
under appropriate integrability assumptions. Variance provides a familiar instance:
[ \operatorname{Var}(X) =\mathbb E[X^2]-(\mathbb E[X])^2\ge0. ] (ocw.mit.edu)
Conditional expectation
Conditional expectation averages a random variable using information supplied by an event or another variable. For an event (A) with positive probability, (\mathbb E[X\mid A]) is computed using the conditional distribution given (A). The quantity (\mathbb E[X\mid Y]), before a particular value of (Y) is specified, is itself a random variable. (ocw.mit.edu)
The law of total expectation states, for integrable (X),
[ \mathbb E[\mathbb E[X\mid Y]]=\mathbb E[X]. ]
For a finite or countable partition into events (A_i) of positive probability, this becomes
[ \mathbb E[X] =\sum_i P(A_i)\mathbb E[X\mid A_i]. ]
It allows a complicated average to be calculated by first averaging within separate cases and then weighting those case averages. (ocw.mit.edu)
Statistical and computational applications
In machine learning, expected loss measures a predictor’s performance over a data-generating distribution. For a loss function (L), a predictor (h) has population risk
[ R(h)=\mathbb E[L(h(X),Y)]. ]
Empirical risk minimization replaces this expectation with an average over training data. The distinction between that observed average and population risk is central to generalization: good performance on a sample does not automatically establish equally good performance on unseen observations. (cs229.stanford.edu)
In reinforcement learning, a value function commonly represents the expected sum of discounted rewards obtained by following a policy from a specified state. The Bellman equation expresses this expectation recursively as immediate reward plus the discounted expected value of the next state. These expectations average over possible trajectories, not just a single realized sequence of actions and outcomes. (cs229.stanford.edu)