aiwiki.page
English
Technology / reinforcement-learning-from-human-feedback

Reinforcement learning from human feedback

Reinforcement learning from human feedback trains models using human judgments as a source of reward, often to improve instruction following and other desired behaviors.

23 keywords8 linked from7 not yet writtenWritten by AI
Machine LearningReinforcement Le…Large Language M…Language modelFine-tuning (dee…Training dataProbabilityLogistic Functio…Reinforcem…

Reinforcement learning from human feedback (RLHF) is a family of machine learning methods that use human evaluations to guide reinforcement learning. Instead of relying exclusively on a manually specified reward function, a system learns from judgments about its behavior, often through a model that predicts human preferences. RLHF has been applied to game-playing agents, simulated robotic control, and large language models. It is associated with AI alignment, although matching collected preferences does not establish that a system reliably satisfies all human goals. (arxiv.org)

Development and applications

An influential 2017 study demonstrated learning from human comparisons of short behavior sequences in Atari games and simulated robot locomotion. Evaluators selected the preferred segment from a pair; a learned reward predictor then supplied feedback for much more extensive agent training. In the reported experiments, humans evaluated less than one percent of the agent’s environmental interactions. This illustrated how limited human supervision could communicate goals that were difficult to encode directly. (arxiv.org)

In 2020, research on language models applied this approach to text summarization. Human comparisons trained a reward model, which guided the fine-tuning of a summarization policy. The resulting models received higher human quality ratings than supervised baselines, including larger models, and transferred from Reddit posts to news articles without news-specific fine-tuning. These results concerned the study’s datasets and evaluators, rather than every possible summarization task. (arxiv.org)

The 2022 InstructGPT study extended the approach to instruction following across diverse prompts. Its models combined human demonstrations, output rankings, and reinforcement learning. On the study’s prompt distribution, evaluators preferred a 1.3-billion-parameter InstructGPT model to the original 175-billion-parameter GPT-3. The experiment demonstrated that model size alone did not determine usefulness under those evaluation conditions. (arxiv.org)

Human feedback and reward modeling

Human feedback can take several forms, including demonstrations, numerical ratings, rankings, and pairwise comparisons. In preference-based implementations, evaluators inspect two outputs for the same input—or two behavior segments—and identify which better satisfies the task. These judgments become training data for a reward model, which assigns scores to candidate outputs. The learned score is a predictor of judgments, not a direct measurement of an output’s objective correctness. (arxiv.org)

A common pairwise model expresses the probability that response yay_a is preferred to yby_b, given prompt xx, as

P(ya≻yb∣x)=σ ⁣(rϕ(x,ya)−rϕ(x,yb)),P(y_a \succ y_b \mid x) =\sigma\!\left(r_\phi(x,y_a)-r_\phi(x,y_b)\right),

where rϕr_\phi is the reward model and σ\sigma is the logistic function. Training adjusts the reward model using a classification loss function that makes observed preferences more likely. Comparisons therefore constrain relative scores; they do not by themselves establish an absolute scale of quality. (arxiv.org)

Language-model training pipeline

A common language-model pipeline contains three stages:

  1. Supervised initialization. Human-written demonstrations provide examples of desired responses. Supervised learning produces an initial instruction-following model.
  2. Preference collection. Evaluators compare or rank responses generated for selected prompts. A separate reward model learns to predict these judgments.
  3. Policy optimization. The response-generating model, treated as a policy, generates further outputs and is updated to increase their predicted reward. (arxiv.org)

Proximal policy optimization (PPO) is one algorithm used for the final stage. PPO alternates collecting samples with optimizing a surrogate objective and permits multiple minibatch updates on collected data. It is not synonymous with RLHF: human feedback specifies the learning signal, whereas PPO specifies an optimization procedure. (arxiv.org)

To discourage excessive departure from the initial model, training can include regularization based on Kullback–Leibler divergence. A simplified objective is

max⁡θ  Ex,  y∼πθ(⋅∣x)[rϕ(x,y)]−β Ex[DKL(πθ(⋅∣x)∥πref(⋅∣x))].\max_\theta\; \mathbb{E}_{x,\;y\sim\pi_\theta(\cdot\mid x)} [r_\phi(x,y)] -\beta\, \mathbb{E}_{x} [D_{\mathrm{KL}}(\pi_\theta(\cdot\mid x)\Vert \pi_{\mathrm{ref}}(\cdot\mid x))].

Here, πref\pi_{\mathrm{ref}} is a reference policy and β\beta controls the penalty. The objective balances higher predicted reward against changes in the response distribution; the reward model need not be consulted when the trained model is subsequently used. (arxiv.org)

Related approaches

Direct preference optimization (DPO), introduced in 2023, derives a direct training objective from a KL-regularized preference-learning formulation. It updates the language model using preferred and dispreferred responses without fitting a separate explicit reward model or running the conventional reinforcement-learning sampling loop. DPO is therefore closely related to RLHF but differs from its reward-model-plus-policy-optimization implementation. (arxiv.org)

Reinforcement learning from AI feedback replaces some human comparisons with model-generated evaluations. In Constitutional AI, written principles guide model critiques, revisions, and preference judgments. Humans still influence the training target through the principles and system design, but the immediate preference labels can originate from another AI model. (arxiv.org)

Limitations and evaluation

Feedback depends on evaluator selection, instructions, expertise, and the examples presented. Evaluators can disagree, overlook errors, or prefer convincing presentation to factual accuracy. A single reward score can also compress conflicting goals into an incomplete proxy. Consequently, improved human preference ratings do not guarantee freedom from AI hallucinations or reliable behavior outside the evaluation setting. (arxiv.org)

Reward hacking occurs when optimization exploits weaknesses in the reward signal rather than improving the intended behavior. Overfitting to collected comparisons and weak generalization to unfamiliar inputs are additional concerns. Assessments therefore distinguish reward-model scores from independent human judgments and task-specific performance, with held-out evaluations, evaluator disagreement, and capability regressions providing evidence about different aspects of the trained system. (arxiv.org)