Jensen’s inequality is a theorem relating convex functions to weighted averages and expected values. It states that applying a convex function after averaging gives a result no greater than averaging after applying the function. For concave functions, the inequality reverses. Its finite and probabilistic forms express the same underlying principle. (web.stanford.edu)
Finite weighted form
Let (C) be a convex set in a real vector space, and let (f:C\to\mathbb R) be convex. For points (x_1,\ldots,x_n\in C) and weights (\lambda_i\geq0) satisfying (\sum_i\lambda_i=1), Jensen’s inequality states
[ f!\left(\sum_{i=1}^{n}\lambda_i x_i\right) \leq \sum_{i=1}^{n}\lambda_i f(x_i). ]
The argument on the left is a convex combination, which belongs to (C) because the domain is convex. Equal weights give the familiar comparison between the function of an arithmetic average and the arithmetic average of the function values. The statement applies to scalar or vector inputs, while the output remains scalar. (stanford.edu)
For two points, this is precisely the defining condition for convexity:
[ f(tx+(1-t)y)\leq tf(x)+(1-t)f(y), \qquad 0\leq t\leq1. ]
Geometrically, the graph lies below the chord joining two graph points. The finite inequality follows by repeatedly applying this two-point condition, or by mathematical induction. Nonnegative, normalized weights are essential; arbitrary signed coefficients do not give the same theorem. (web.stanford.edu)
Expectation and integral forms
Let (X) be an integrable random variable taking values in an open interval (I), and let (f:I\to\mathbb R) be convex. A standard finite-valued formulation assumes that (f(X)) is also integrable. Then
[ f(\mathbb E[X])\leq\mathbb E[f(X)]. ]
The finite weighted version is recovered when the probability distribution assigns mass (\lambda_i) to (x_i). No assumption of independence is involved. The distinction is between transforming a mean and taking the mean of a transformation; these operations generally do not commute. (mit.edu)
In measure-theoretic notation, for a probability space ((\Omega,\mathcal F,\mu)),
[ f!\left(\int_\Omega X,d\mu\right) \leq \int_\Omega f(X),d\mu. ]
These integrals use a measure of total mass one. For a finite positive measure of mass (M>0), normalization introduces (1/M) on both sides. Integrability hypotheses matter: an undefined mean cannot simply be substituted into the formula. (stanford.edu)
There is also a conditional version. Under suitable integrability assumptions, for a sub-sigma-algebra (\mathcal G),
[ f(\mathbb E[X\mid\mathcal G]) \leq \mathbb E[f(X)\mid\mathcal G] \quad\text{almost surely}. ]
Thus the same comparison holds for conditional expectations, with the inequality interpreted almost surely rather than necessarily at every outcome. (dspace.mit.edu)
Proof and equality conditions
A useful proof uses a supporting line. Write (m=\mathbb E[X]). If (f) is differentiable and convex on an open interval containing (m), its derivative gives
[ f(x)\geq f(m)+f'(m)(x-m). ]
Taking expectations makes the last term vanish, because (\mathbb E[X-m]=0), leaving Jensen’s inequality. The proof also explains its direction: convexity places the function above its tangent line. (live.ocw.mit.edu)
Differentiability is not essential. At an interior point of a finite convex function’s domain, an appropriate supporting slope can replace the derivative. In several dimensions, the analogous argument uses a supporting hyperplane and, for differentiable functions, the gradient. This connects Jensen’s inequality with the first-order characterization of convexity used in convex optimization. (web.stanford.edu)
If (f) is strictly convex, equality in the finite form holds exactly when all points with positive weight coincide. In the probabilistic form, equality holds exactly when (X) is constant almost surely. For a non-strictly convex function, equality can also occur across an interval on which the function is affine. An affine function gives equality for every admissible distribution. (cs229.stanford.edu)
Examples and the Jensen gap
Taking (f(x)=x^2) gives, for a variable with finite second moment,
[ (\mathbb E[X])^2\leq\mathbb E[X^2]. ]
The difference is the variance of (X). As a concrete two-point example, averaging (1) and (3) before squaring gives (4), whereas averaging their squares gives (5). This illustrates the effect without requiring any probability notation. (mit.edu)
Because the logarithm is concave, positive numbers satisfy
[ \sum_i\lambda_i\log x_i \leq \log!\left(\sum_i\lambda_i x_i\right). ]
Exponentiation yields the weighted arithmetic–geometric mean inequality. Conversely, the convex exponential function gives (e^{\mathbb E[X]}\leq\mathbb E[e^X]) whenever the expectations are defined. (cs.cmu.edu)
The nonnegative difference
[ J_f(X)=\mathbb E[f(X)]-f(\mathbb E[X]) ]
is called the Jensen gap. Beyond its sign, bounds on its magnitude can depend on the function’s growth or curvature and on moments describing the distribution’s dispersion. Such bounds distinguish a nearly tight inequality from one with a substantial gap. (arxiv.org)
Applications and historical origin
In machine learning, logarithmic Jensen inequalities construct lower bounds on otherwise difficult likelihood expressions. For a discrete latent variable (z) and a positive auxiliary distribution (q),
[ \log p(x)
\log\mathbb E_q!\left[\frac{p(x,z)}{q(z)}\right] \geq \mathbb E_q!\left[\log\frac{p(x,z)}{q(z)}\right]. ]
With appropriate support assumptions, the right side is an evidence lower bound. This construction underlies variational inference and the expectation–maximization algorithm, where choosing the current conditional distribution of the latent variable makes the bound tight. (cs229.stanford.edu)
The inequality bears the name of Johan Ludwig William Valdemar Jensen. His 1906 paper, Sur les fonctions convexes et les inégalités entre les valeurs moyennes, developed convex-function inequalities and their relationships with classical inequalities between means. (cs.cmu.edu)