A random variable is a function that assigns a numerical value to each possible outcome of a random experiment. It connects outcomes, such as sequences of coin tosses, with quantities, such as the number of heads. In probability theory, a real-valued random variable is formally a measurable function on a probability space. Its defining rule is fixed; uncertainty concerns which outcome occurs and therefore which value is observed. (ocw.mit.edu)
Mathematical definition
A probability space is written ((\Omega,\mathcal F,P)). Here, (\Omega) is the sample space of possible outcomes, (\mathcal F) is a sigma-algebra of events, and (P) assigns probabilities to those events. A real-valued random variable is a function [ X:\Omega\longrightarrow\mathbb R ] such that [ {\omega\in\Omega:X(\omega)\leq x}\in\mathcal F ] for every real number (x). This measurability condition ensures that statements about the value of (X) correspond to events with defined probabilities. It belongs to the framework of measure theory. (math.mit.edu)
An observed value (X(\omega)) is called a realization. Conventionally, capital letters denote random variables and lowercase letters denote their possible values. Different outcomes can produce the same realization: when two dice are rolled, the outcomes ((1,6)) and ((3,4)) both give (X=7) if (X) records their sum. A constant function is also a random variable, although its value has no uncertainty. (ocw.mit.edu)
Distribution and classification
The probability distribution, or law, of (X) describes how probabilities are assigned to its values. For a Borel set (B), [ P_X(B)=P{\omega:X(\omega)\in B}. ] Thus, a random variable is a function on outcomes, whereas its distribution is a probability measure on values. Distinct random variables can have identical distributions without being equal outcome by outcome. (math.mit.edu)
The cumulative distribution function (CDF) is [ F_X(x)=P(X\leq x). ] It completely determines the law. Every CDF is nondecreasing and right-continuous, with limits (0) and (1) as (x) tends to negative and positive infinity. Interval probabilities follow from [ P(a<X\leq b)=F_X(b)-F_X(a). ] (math.mit.edu)
A discrete random variable takes values in a finite or countably infinite set, apart from possible outcomes of probability zero. Its probability mass function is (p_X(x)=P(X=x)), with nonnegative masses summing to one. A success indicator has a Bernoulli distribution; a count of arrivals may be modeled by a Poisson distribution. (ocw.mit.edu)
An absolutely continuous random variable has a probability density function (f_X), satisfying [ P(a<X\leq b)=\int_a^b f_X(x),dx. ] The density integrates to one, but its height is not a probability and can exceed one. Every individual value has probability zero. The normal distribution is an important example. Not all laws are discrete or absolutely continuous: mixed laws combine point masses with a continuous component, while singular continuous laws have continuous CDFs but no density with respect to ordinary length measure. (ocw.mit.edu)
Expectation and variability
The expected value summarizes a distribution through a probability-weighted average. For discrete and absolutely continuous variables, respectively, [ E[X]=\sum_x xp_X(x), \qquad E[X]=\int_{-\infty}^{\infty}xf_X(x),dx. ] A finite expectation requires (E[|X|]<\infty); not every random variable satisfies this condition. An expectation need not be a possible realization. For example, a fair six-sided die has mean (3.5). (ocw.mit.edu)
When the second moment is finite, variance measures dispersion: [ \operatorname{Var}(X) =E[(X-E[X])^2] =E[X^2]-(E[X])^2. ] Its square root is the standard deviation. For constants (a,b), [ E[aX+b]=aE[X]+b,\qquad \operatorname{Var}(aX+b)=a^2\operatorname{Var}(X). ] Expectation is linear even for dependent variables; additivity of variance requires additional conditions, such as independence. (ocw.mit.edu)
Transformations and dependence
If (g) is a measurable function, then (Y=g(X)) is another random variable. Its law is obtained by identifying the original values mapped into each target set. In general, (E[g(X)]) is not (g(E[X])): squaring a variable, for example, involves its second moment rather than merely its mean. (ocw.mit.edu)
Several variables defined on the same probability space have a joint probability distribution. They are independent when [ P(X\in A,Y\in B)=P(X\in A)P(Y\in B) ] for all Borel sets (A,B). Knowing their individual distributions alone does not establish independence. Covariance measures one form of association, but zero covariance does not generally imply independence. Information about one variable can instead be represented through conditional distributions and conditional expectation. (ocw.mit.edu)
Sequences and statistical use
A stochastic process is an indexed family of random variables, often representing quantities evolving over time. For independent, identically distributed variables with finite absolute expectation, the law of large numbers establishes convergence of their sample average to the common mean. With finite, positive variance, the classical central limit theorem describes convergence of the standardized average’s distribution to the standard normal law. These results concern different aspects of averaging: stabilization and the distribution of fluctuations. (ocw.mit.edu)
In statistics, observations are modeled as realizations of random variables. A statistic, being a function of the observations, is itself random before the data are observed. Its sampling distribution describes how its value varies across hypothetical repeated samples and provides a basis for quantifying estimation uncertainty. (ocw.mit.edu)