In reinforcement learning, a policy specifies how an agent selects actions while interacting with an environment. It maps the information available to the agent to an action or a distribution over actions. A policy describes behavior, rather than the environment’s dynamics or the rewards assigned to outcomes. Learning a policy that maximizes long-term expected reward is a central objective of reinforcement learning, a branch of machine learning. Policies may be explicitly represented and optimized, or implicitly determined by action-value estimates and an action-selection rule. (spinningup.openai.com)
Mathematical definition
In a fully observable Markov decision process (MDP), a stationary stochastic policy is commonly written
For each state , it defines a probability distribution over available actions. With discrete actions, probabilities are nonnegative and sum to one. A deterministic policy instead selects a single action, often written . Determinism concerns the agent’s selection rule: a deterministic policy can still operate in an environment with random transitions. (spinningup.openai.com)
A stationary policy uses the same rule at every decision time. A nonstationary policy may depend explicitly on time, as . This distinction matters in finite-horizon problems, where the best action can change as the remaining time decreases. Stationarity does not mean that a learning algorithm leaves its policy unchanged during training; it describes the absence of explicit time dependence in a particular policy. (hankyang.seas.harvard.edu)
An MDP state satisfies the Markov property, making it sufficient for predicting subsequent transitions given an action. In a partially observable problem, the agent receives observations rather than the complete state. Policies can therefore use observation histories or internal memory, including representations maintained by a recurrent neural network. Reacting only to the latest observation can discard information relevant to decisions. (hankyang.seas.harvard.edu)
Return, value, and optimality
Policies are evaluated through their consequences over sequences of interactions. A common objective is the expected value of discounted return:
where is the initial-state distribution. The discount factor controls how strongly later rewards contribute. Finite-horizon total reward is another objective. Maximizing immediate reward alone need not maximize return, because actions influence future states and opportunities. (spinningup.openai.com)
A value function estimates return under a policy. The state value assumes that the agent starts in and follows ; the action value assumes that it first takes and follows thereafter. These quantities satisfy a Bellman equation, connecting immediate reward to subsequent value. For discrete actions,
Thus, a policy selects actions, whereas a value function assesses their expected consequences. (spinningup.openai.com)
For a finite MDP with bounded rewards and infinite-horizon discounted return, an optimal deterministic stationary policy exists. If the optimal action values are known, it can choose any maximizing action:
Several actions may tie. This existence result concerns the stated MDP setting, not every constrained, partially observable, or otherwise modified decision problem. (hankyang.seas.harvard.edu)
Representation and learning
Small policies can be represented as tables of actions or action probabilities. Larger problems use function approximation, including an artificial neural network parameterized by . For discrete actions, a softmax function can convert network outputs into probabilities. For continuous actions, a policy may output an action directly or specify a distribution, such as a Gaussian distribution whose parameters depend on the state. (spinningup.openai.com)
Several algorithmic approaches connect policy representation with learning:
- Value-based learning: Q-learning estimates action values; a selection rule derives behavior from those estimates.
- Policy optimization: a policy-gradient method directly adjusts policy parameters to increase expected return.
- Combined learning: an actor–critic method learns both an action-selecting actor and a value-estimating critic. (spinningup.openai.com)
Policy-gradient updates use the gradient of action log probabilities, weighted by estimates of return or advantage. Advantage measures an action’s value relative to the policy’s state value. Value-based baselines can reduce estimator variance without changing the expected policy gradient under the relevant assumptions. Proximal policy optimization uses a surrogate objective; its clipped variant removes incentives for certain excessively large probability-ratio changes, rather than imposing a strict bound on every policy change. (spinningup.openai.com)
Exploration and data collection
During learning, action selection must address the exploration–exploitation trade-off: exploiting currently promising actions versus trying alternatives that may reveal better behavior. An epsilon-greedy policy usually selects a greedy action but occasionally samples an exploratory action. Stochastic policies explore by sampling from their action distributions; deterministic policies can collect exploratory experience by adding action noise. Randomness alone does not guarantee adequate exploration. (hankyang.seas.harvard.edu)
The behavior policy generates training interactions, while the target policy is evaluated or improved. In off-policy learning, these policies may differ. On-policy methods instead learn using interactions generated by the policy being evaluated or improved, subject to the algorithm’s particular update procedure. This distinction concerns the relationship between data collection and learning—not whether the policy is deterministic or stochastic. For example, deterministic actor–critic algorithms can learn off-policy while their behavior policy adds exploration noise. (spinningup.openai.com)