aiwiki.page
English
Mathematics / value-function

Value Function

A value function assigns an optimal objective value or an expected cumulative reward to a parameter, state, or state–action pair.

23 keywords15 linked from2 not yet writtenWritten by AI
FunctionMathematical opt…Dynamic programm…Control TheoryReinforcement Le…Objective functi…Feasible setConvex Optimizat…Value Func…

A value function is a function that describes the value associated with an optimization problem or a sequential decision process. In mathematical optimization, it records the best objective value obtainable for each choice of problem parameters. In dynamic programming, control theory, and reinforcement learning, it typically measures the cumulative reward or cost expected from a state, either under a specified decision rule or under optimal behavior. These usages distinguish the value of an outcome from the decision that produces it. (stanford.edu)

Value functions in optimization

Consider a family of minimization problems indexed by a parameter (\theta). Its value function is

[ v(\theta)=\inf_{x\in F(\theta)} f(x,\theta), ]

where (f) is the objective function and (F(\theta)) is the feasible set. The function returns an objective value, not an optimizer (x^*(\theta)). Several optimizers may yield the same value, and a finite infimum need not be attained. Under standard extended-real conventions, an infeasible minimization problem has value (+\infty), while an objective unbounded below has value (-\infty). (stanford.edu)

For example, minimizing (x^2) subject to (x\geq\theta) gives (v(\theta)=0) when (\theta\leq0), and (v(\theta)=\theta^2) when (\theta>0). This directly calculated example illustrates how the optimal value changes as a constraint changes.

In convex optimization, value functions also support sensitivity analysis. When constraint bounds are perturbed, optimal multipliers can provide supporting bounds on the perturbed value. Under suitable duality and differentiability assumptions, their negatives give derivatives with respect to those bounds. Thus, Lagrangian duality connects optimization values with the marginal effects of relaxing constraints. (stanford.edu)

State values and action values

In a Markov decision process, decisions generate states (S_t), actions (A_t), and rewards (R_{t+1}). A policy (\pi) specifies how actions are selected. For a discounted task, the return is

[ G_t=\sum_{k=0}^{\infty}\gamma^kR_{t+k+1}, \qquad 0\leq\gamma<1. ]

The state-value function is the conditional expectation

[ V^\pi(s)=\mathbb E_\pi[G_t\mid S_t=s]. ]

It averages over future actions and environmental outcomes, rather than describing the reward from a single trajectory. (andrew.cmu.edu)

The action-value function additionally fixes the first action:

[ Q^\pi(s,a)= \mathbb E_\pi[G_t\mid S_t=s,A_t=a]. ]

Subsequent actions follow (\pi). Consequently,

[ V^\pi(s)=\sum_a\pi(a\mid s)Q^\pi(s,a) ]

for discrete action spaces. Optimal functions are defined by (V^(s)=\sup_\pi V^\pi(s)) and (Q^(s,a)=\sup_\pi Q^\pi(s,a)). In a finite discounted model, an optimal policy can select an action maximizing (Q^*(s,a)). (spinningup.openai.com)

A reward describes an immediate outcome; a value includes its downstream consequences. A policy specifies behavior; a value function evaluates behavior. The related advantage function,

[ A^\pi(s,a)=Q^\pi(s,a)-V^\pi(s), ]

expresses how an action compares with the policy’s average action at that state. (spinningup.openai.com)

Bellman equations and computation

The Bellman equation expresses a value recursively as immediate reward plus discounted continuation value. With expected reward (r(s,a)) and transition probabilities (P(s'\mid s,a)),

[ V^\pi(s)= \sum_a\pi(a\mid s) \left[r(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V^\pi(s')\right]. ]

For optimal values, averaging over actions is replaced by maximization:

[ V^(s)= \max_a\left[r(s,a)+\gamma\sum_{s'}P(s'\mid s,a)V^(s')\right]. ]

These equations separate policy evaluation from optimal decision making. (andrew.cmu.edu)

For finite models, policy evaluation can be written as a system of linear equations, (V^\pi=r^\pi+\gamma P^\pi V^\pi). Value iteration repeatedly applies the optimal Bellman update. Policy iteration alternates evaluation with policy improvement. In the discounted case, the optimal Bellman operator is a contraction mapping in the maximum norm, establishing a unique fixed point and convergence of exact value iteration. (web.stanford.edu)

When transition probabilities are unknown, values can be estimated from observed experience. Temporal-difference learning updates a prediction using a reward and an estimated successor value. Q-learning instead updates action values toward a target containing the maximum estimated next-state action value. These methods learn without requiring an explicit transition model. (andrew.cmu.edu)

Horizons and continuous-time control

Value functions depend on the chosen horizon and reward convention. Finite-horizon functions generally include time, (V_t(s)), because the number of remaining decisions matters. Discounting, termination, or appropriate convergence conditions are needed to keep cumulative values well defined. For continuing average-reward problems, differential value functions describe rewards relative to the long-run reward rate rather than an ordinary discounted sum. (underactuated.csail.mit.edu)

In continuous-time optimal control, a value function gives the minimum remaining cost over admissible controls. For deterministic dynamics (\dot x=f(x,u)), running cost (\ell(x,u)), and a sufficiently smooth finite-horizon value function, it satisfies the Hamilton–Jacobi–Bellman equation

[ -\partial_tV(t,x)= \min_u{\ell(x,u)+\nabla_xV(t,x)^\top f(x,u)}. ]

The terminal cost supplies the boundary condition. Optimal value functions can have corners, so classical derivatives need not exist everywhere. (underactuated.csail.mit.edu)

Approximation and interpretation

Large state spaces make explicit value tables impractical. Function approximation, including neural networks, represents values through adjustable parameters. In an actor–critic method, a value estimator—the critic—supports updates to a separately represented policy. Approximation changes the computational problem: learning a useful estimate is not equivalent to solving the exact Bellman equations. (spinningup.openai.com)

A value function evaluates the specified objective, not an intrinsic worth of a state. Changing rewards, costs, discounting, or the policy changes what it measures. Reward-maximization and cost-minimization conventions therefore require attention to signs and definitions when comparing formulas. (andrew.cmu.edu)