A probability mass function (PMF) is a function that assigns to each possible value of a discrete random variable the probability of that value occurring. It completely specifies a discrete probability distribution, whose probability is concentrated on a finite or countably infinite collection of values. Common notations include (p_X(x)), (p(x)), and (f(x)); the subscript identifies the variable when several distributions are under consideration. (online.stat.psu.edu)
Definition and basic properties
For a discrete random variable (X), its PMF is defined by
[ p_X(x)=P(X=x). ]
Let (S={x:p_X(x)>0}) denote the set of values with positive probability, commonly called its support in elementary treatments. This is a countable set, and the PMF satisfies
[ p_X(x)\geq 0,\qquad \sum_{x\in S}p_X(x)=1. ]
Values outside (S) have probability zero. These requirements are also sufficient: any nonnegative assignment to a finite or countable set that sums to one defines a discrete distribution. For any subset (A) of possible values,
[ P(X\in A)=\sum_{x\in A\cap S}p_X(x). ]
Thus probabilities of compound events are obtained by adding the masses of their constituent values. (online.stat.psu.edu)
“Discrete” does not mean that the values must be integers or equally spaced. The essential requirement is that all probability lies on a countable collection of values. A PMF can be represented by a formula, a table, or a graph with separate markers or bars. The heights represent individual probabilities, rather than a continuous density. (online.stat.psu.edu)
Examples and common distributions
For a fair six-sided die, let (X) be the number on the upper face. Then
[ p_X(k)= \begin{cases} 1/6,&k\in{1,2,3,4,5,6},\ 0,&\text{otherwise}. \end{cases} ]
Consequently, (P(X>4)=p_X(5)+p_X(6)=1/3). This illustrates a finite distribution with equal masses, whereas many discrete distributions assign different probabilities to different values. (online.stat.psu.edu)
The Bernoulli distribution models a single binary outcome. With success probability (q), it assigns mass (q) to 1 and (1-q) to 0. The binomial distribution describes the number of successes in (n) independent Bernoulli trials with the same success probability:
[ p_X(k)=\binom nk q^k(1-q)^{n-k}, \qquad k=0,\ldots,n. ]
The binomial coefficient accounts for the different arrangements of (k) successes among the trials. (online.stat.psu.edu)
An important infinite-support example is the Poisson distribution, with parameter (\lambda>0):
[ p_X(k)=e^{-\lambda}\frac{\lambda^k}{k!}, \qquad k=0,1,2,\ldots. ]
It assigns probability to every nonnegative integer and is used in models of event counts. Although infinitely many terms occur, their sum is one. (online.stat.psu.edu)
Relationship to cumulative distribution and density
For a real-valued discrete variable, the cumulative distribution function (CDF) collects all masses at or below a threshold:
[ F_X(t)=P(X\leq t)= \sum_{\substack{x\in S\x\leq t}}p_X(x). ]
At a value (x) with positive mass, the CDF has a jump of size (p_X(x)). Equivalently,
[ p_X(x)=F_X(x)-F_X(x^-), ]
where (F_X(x^-)) is the left-hand limit. For integer-valued variables, this becomes (p_X(k)=F_X(k)-F_X(k-1)). Unlike the PMF, the CDF is defined through the same probability expression for both discrete and continuous variables. (online.stat.psu.edu)
A PMF must be distinguished from a probability density function (PDF). A PMF value is itself a probability and cannot exceed one. For a distribution described by a continuous density, interval probabilities are obtained through an integral, and the probability of any single value is zero. Density values are not individual probabilities and may exceed one; their total integral, rather than their sum, equals one. (online.stat.psu.edu)
Expectations and transformations
The PMF supplies the weights used to calculate the expected value:
[ E[X]=\sum_{x\in S}x,p_X(x), ]
provided the series is absolutely convergent when a finite expectation is intended. More generally,
[ E[g(X)]=\sum_{x\in S}g(x)p_X(x). ]
When the relevant moments are finite, the variance is
[ \operatorname{Var}(X) =\sum_{x\in S}(x-\mu)^2p_X(x) =E[X^2]-\mu^2, \qquad \mu=E[X]. ]
Its square root is the standard deviation. A valid PMF need not have a finite mean or variance, particularly when its support is infinite. (online.stat.psu.edu)
Joint and conditional mass functions
For two discrete variables, a joint probability distribution is specified by
[ p_{X,Y}(x,y)=P(X=x,Y=y). ]
Summing over the other variable gives an individual, or marginal, PMF:
[ p_X(x)=\sum_y p_{X,Y}(x,y). ]
By the definition of conditional probability, when (p_Y(y)>0),
[ p_{X\mid Y}(x\mid y) =\frac{p_{X,Y}(x,y)}{p_Y(y)}. ]
Statistical independence holds exactly when the joint PMF factorizes as (p_{X,Y}(x,y)=p_X(x)p_Y(y)) for all pairs of values. (online.stat.psu.edu)
Role in statistical inference
In statistics, a parameterized PMF (p(x;\theta)) describes how probabilities depend on an unknown parameter. For independent observations (x_1,\ldots,x_n) sharing that PMF, their product defines the likelihood function:
[ L(\theta)=\prod_{i=1}^{n}p(x_i;\theta). ]
Maximum likelihood estimation selects parameter values that maximize this expression. The distinction is interpretive: a PMF treats the outcome as variable with the parameter fixed; a likelihood treats the observed data as fixed and compares parameter values. A likelihood is therefore not generally a probability distribution over the parameter. (online.stat.psu.edu)