Reinforcement learning (RL) is a branch of machine learning concerned with learning how to act through experience. An agent selects actions in an environment, receives numerical rewards, and seeks a decision-making rule that maximizes cumulative reward. Unlike supervised learning, it generally receives evaluative feedback rather than a correct action for every situation; unlike unsupervised learning, its defining objective is successful behavior rather than discovering structure in data. Actions can influence both future outcomes and the experience available for learning. (web.stanford.edu)
Agent–environment framework
At each time step, an agent observes information about its situation, chooses an action, and receives a reward and a subsequent observation. A policy specifies how actions are selected: it may prescribe one action deterministically or assign probabilities to several possibilities. A task can be episodic, ending at a terminal condition, or continuing without a natural endpoint. Rewards evaluate outcomes but need not identify which preceding decisions produced them. (web.stanford.edu)
The standard mathematical framework is a Markov decision process, comprising states, actions, transition probabilities, rewards, and an objective. Its Markov property means that the current state contains the information needed to characterize the next-state and reward distribution, conditional on the action. When observations do not reveal this state, a partially observable Markov decision process provides a more general formulation; the agent may need memory or an estimate of hidden state. (web.stanford.edu)
A common objective maximizes the expected value of the discounted return:
Here, is a future reward and discounts more distant outcomes. With bounded rewards, this infinite sum is finite. Other objectives include undiscounted finite-episode returns and long-run average reward. A value function estimates expected return from a state, or from a state–action pair, under a specified policy. (spinningup.openai.com)
Learning and exploration
Two central difficulties are exploration and credit assignment. The exploration–exploitation trade-off concerns whether to try uncertain actions or choose actions already believed to be rewarding. An epsilon-greedy policy, for example, usually selects the estimated best action but sometimes chooses randomly. Credit assignment concerns identifying which decisions contributed to a later outcome, particularly when rewards are delayed or infrequent. (web.stanford.edu)
Value estimates exploit recursive relationships expressed by the Bellman equation, connecting immediate rewards to future values. If the environment’s dynamics are known, dynamic programming methods can compute policies using these relationships. Learning methods instead estimate relevant quantities from sampled experience. (web.stanford.edu)
Monte Carlo methods estimate values from completed returns. Temporal-difference learning updates predictions using rewards and other predictions, allowing learning before an episode ends. Q-learning estimates optimal action values using a target based on the highest estimated next-state action value. SARSA instead uses the action actually selected next. This illustrates the distinction between on-policy learning and off-policy learning, which learns about a policy different from the one generating experience. (web.stanford.edu)
Algorithm families
Model-based methods use a known or learned model of transitions and rewards to plan or generate simulated experience. Model-free methods learn policies or values without explicitly using such a model for planning. Models can improve sample efficiency, but inaccurate predictions may lead to decisions that succeed in simulation and fail in the actual environment. These categories concern the use of environmental models, not whether the agent uses a neural network. (spinningup.openai.com)
Value-based methods derive actions from estimated values. Policy-gradient methods directly adjust policy parameters to improve expected return. Actor–critic methods combine an actor that selects actions with a critic that estimates values and informs policy updates. Proximal policy optimization is a policy-optimization method designed to limit disruptive updates through a constrained or clipped surrogate objective. These approaches can accommodate discrete or continuous actions, depending on their particular formulation. (spinningup.openai.com)
Deep and offline reinforcement learning
Deep reinforcement learning combines RL with deep learning, using an artificial neural network to approximate policies, values, or environmental models. The deep Q-network demonstrated learning from image inputs across 49 Atari games in a 2015 study. Its stabilizing mechanisms included experience replay, which reuses stored transitions, and a separate target network that changes more slowly than the network being trained. (nature.com)
Offline reinforcement learning learns from a fixed collection of previously recorded interactions without additional environmental data collection. It differs from ordinary off-policy learning, which may continue gathering experience. A major difficulty is distribution shift: the learned policy may select actions insufficiently represented in the dataset, producing unreliable value estimates. Policy constraints and conservative value estimation are among the approaches developed to address this problem. (arxiv.org)
Applications and limitations
Applications include game playing and robotics. AlphaGo combined learning from expert games, reinforcement learning through self-play, and tree search. In language-model training, reinforcement learning from human feedback can use human comparisons to train a reward model, whose predictions then guide policy optimization rather than serving as direct demonstrations of every desired output. (research.google)
Practical limitations include costly interaction, unsafe exploration, inaccurate models, and failure under unfamiliar conditions. A reward is an operational objective, not a complete specification of intended behavior. Reward hacking occurs when an agent exploits weaknesses in that objective to obtain high reward while violating the designer’s intent. Consequently, high measured return alone does not establish safety or successful performance outside the evaluated environment. (spinningup.openai.com)