RLHF (Reinforcement Learning from Human Feedback) and LLM Alignment
1. Overview
RLHF (Reinforcement Learning from Human Feedback) is a post-training technique that trains a reward model on human-ranked preference data and then performs reinforcement learning using that reward as a signal, so that a large language model (LLM) produces outputs aligned with human intent, values, and safety standards.
An LLM that has only completed pre-training is merely a model optimized to mimic next-token probabilities over a vast web corpus, so it systematically diverges from "fluent output that is actually what the user wants." This mismatch appears in three forms. First, failure to follow instructions: asked to "summarize in three sentences," it pours out a long passage—generating a probabilistically plausible continuation rather than treating the instruction as a task. Second, hallucination and overconfidence: it asserts baseless facts confidently, a result of imitating the "assertive narration" pattern common in corpora. Third, harmfulness: it reproduces harmful patterns—discrimination, violence, assistance with illegal acts—mixed into the training data. RLHF emerged precisely to bridge the gap between the training objective (next-token prediction) and actual utility (helpfulness, honesty, harmlessness, the so-called 3H: Helpful, Honest, Harmless).
The goals that alignment pursues are often summarized as the 3H. Helpful means the ability to grasp the user's real intent and complete the task; Honest means the attitude of admitting what one does not know and not distorting facts; Harmless means the safety of refusing dangerous, discriminatory, or illegal assistance. The problem is that these three frequently conflict. For example, being "fully helpful" to a dangerous question requires sacrificing harmlessness, while leaning too far toward harmlessness undermines helpfulness through over-refusal that rejects even legitimate requests. In that RLHF makes a model learn the balance point among conflicting values from human preference, it has the character of value alignment, beyond a mere performance-improvement technique.
The fundamental reason RLHF is not solved by supervised fine-tuning (SFT) alone is that comparing two options is far easier and cheaper than writing each correct answer out. Having experts write "good answers" from scratch explodes costs, but having them choose which of two responses is better lets even non-experts rapidly produce a large volume of signal. Moreover, goals such as the quality, politeness, and safety of a summary are non-differentiable and hard to quantify, so they are difficult to express directly as a loss function, yet they can be optimized indirectly through a "learned evaluator"—the reward model. When OpenAI's InstructGPT and ChatGPT dramatically raised usability this way in 2022, RLHF became the de facto standard recipe for generative-AI alignment.
RLHF's roots trace back to 2017 research by OpenAI and DeepMind that "learned rewards from human preferences" in robot and game control, later extending through the summarization task (2020) to general-purpose instruction following (2022). That is, RLHF is not a technique that appeared suddenly for LLMs but the result of grafting onto language models a long tradition in reinforcement learning of drawing human judgment into problems where the reward is hard to design directly. Understanding this lineage makes clear that the recent RLVR, which puts "rule-based verification" in the place of "human judgment," is likewise an extension of the same problem awareness.
2. The Three-Stage RLHF Pipeline and Its Components
RLHF is not a single algorithm but a three-stage pipeline running SFT → reward modeling → reinforcement learning. Each stage takes the previous stage's output as input, and the whole can be seen as a process of "distilling human preference into the model's interior."
graph TD
P["Pre-trained LLM<br/>(Base Model)"] --> S["Stage 1: SFT<br/>supervised learning on instruction-response data"]
S --> G["sample multiple responses"]
G --> H["humans compare response pairs by preference<br/>(A vs B ranking)"]
H --> R["Stage 2: train reward model (RM)<br/>predict scalar preference score"]
S --> PO["Stage 3: reinforcement learning (PPO)<br/>update policy LLM"]
R -->|"reward signal"| PO
PO -->|"aligned model"| F["Aligned LLM"]
S -.->|"KL reference model"| PO
Stage 1, supervised fine-tuning (SFT), fine-tunes the base model on high-quality instruction-response demonstration data, granting the basic ability to "follow instructions in conversational form." It is the stage that places the model at the "starting line" with tens of thousands to hundreds of thousands of demonstrations, and it later serves as both the initial policy for reinforcement learning and the reference point that prevents excessive deviation. If SFT is weak, the search space is distorted no matter how refined the later stages are, so in practice the quality of SFT data often determines final performance.
Stage 2, reward model (RM) training, proceeds on data in which humans relatively compared (ranked) multiple responses generated by the policy model for the same prompt. Having humans assign absolute scores from 1 to 10 produces noise because each evaluator's criteria differ, but the pairwise comparison "A is better than B" is highly consistent. The reward model has a structure in which the base LLM's head is replaced with a scalar output, and it is trained based on the Bradley-Terry model to assign a higher score to the preferred response. That is, the loss takes the form of a logistic applied to the score difference between preferred and non-preferred responses, through which it approximates humans' implicit utility function.
Stage 3, reinforcement learning, takes the SFT model as the initial policy and updates the policy using the score given by the reward model as the reward. Here the LLM is formalized as "an agent that generates a token sequence (action) from a prompt (state) and receives a reward," and the most widely used algorithm is PPO (Proximal Policy Optimization). Unlike reinforcement learning in Go or robot control, in language models the action space is ultra-high-dimensional—continuous generation of tens of thousands of tokens—and the reward is given only once for the whole response, which makes it difficult. The core design of this stage is to tie the policy down with a KL-divergence (Kullback-Leibler Divergence) penalty so it does not stray too far from the SFT reference model and chase the reward blindly; the reason is detailed in the next section.
The reason for separating the three stages is that each stage handles a signal of a different nature. SFT handles the absolute demonstration of "what should be written," reward modeling handles the relative preference of "which of the two is better," and reinforcement learning handles the exploratory signal of "how good is a newly generated response." In particular, the reinforcement-learning stage performs online exploration in which the model itself generates responses not in the dataset and receives rewards for them, so it can move the policy even into regions hard to reach with SFT alone, which imitates fixed answers. This is the core answer to the question "isn't SFT enough?"; indeed, even in the InstructGPT experiments the RLHF model consistently received higher preference in human evaluation than the SFT model.
The form of feedback that RLHF handles also needs to be distinguished. The "demonstration" used by SFT is costly to write but the most direct signal; the "comparison" used by the reward model is cheap and highly consistent, favorable for mass collection; and some techniques even use binary "good/bad" feedback or numeric scores. Depending on which form is chosen, the data-collection cost, evaluator consistency, and the applicable training algorithm (PPO, DPO, KTO, etc.) all differ, so designing the feedback form to fit the task and budget becomes the first decision of an alignment project.
The procedure for collecting preference data itself also governs quality. In practice, 4 to 9 responses the policy generated for the same prompt are presented to evaluators to be ranked, then decomposed into many pairwise comparisons for use in training. If the annotation guideline is vague at this point, inter-annotator disagreement grows and the reward model learns noise, so it is important to codify in advance the priority among helpfulness, honesty, and harmlessness and the criteria for handling edge cases.
| Stage | Input | Output | Data nature | Scale (example) |
|---|---|---|---|---|
| Pre-training | web corpus | base model | unsupervised (autoregressive) | trillions of tokens |
| 1. SFT | instruction-response demos | initial policy | supervised (writing answers) | tens of thousands to hundreds of thousands |
| 2. Reward model | response-pair preferences | reward model | pairwise ranking | tens of thousands to millions of pairs |
| 3. RL (PPO) | prompts + RM | aligned model | reinforcement (reward signal) | hundreds of thousands of prompts |
Even after the three stages end, alignment does not stop. User feedback the deployed model receives (likes/dislikes, rewrite requests, reports) is fed back as preference data to improve the next-generation reward model and policy, forming a data flywheel. The better this cycle works in a service, the more alignment fits the actual usage distribution, and this is a major driver by which large commercial services rapidly overtake their initial model quality.
3. How the Reward Model and PPO Optimization Work
The interior of the Stage-3 reinforcement learning is an online loop in which generation → evaluation → update repeats. When the policy generates a response, the reward model scores it, and the policy's parameters are adjusted with a signal combining that reward and a KL penalty. Viewed structurally, four kinds of models interact simultaneously.
graph LR
PR["prompt"] --> POL["policy model<br/>(Policy, trained)"]
POL -->|"generate response"| RM["reward model<br/>(Reward)"]
POL -->|"log-prob"| REF["reference model<br/>(Reference, fixed)"]
RM -->|"reward r"| ADV["reward correction<br/>r - β·KL"]
REF -->|"compute KL"| ADV
ADV --> CRI["value model<br/>(Critic)"]
CRI -->|"advantage"| UPD["PPO update<br/>(clipping)"]
UPD -->|"update parameters"| POL
Reward hacking and the KL constraint. Since the reward model is only an imperfect approximation of human preference, if the policy pursues reward to the extreme it converges to degenerate output that exploits only the reward model's loopholes (e.g., repeating certain clichés, excessive flattery, meaningless long passages)—this is reward hacking. To prevent it, the final objective is designed in the form "reward − β·KL(policy‖reference model)," so the policy is penalized if it drifts too far from the SFT reference model in token distribution. β (the KL coefficient) is the key hyperparameter that adjusts the trade-off between alignment strength and preserving language quality: too small causes reward hacking, too large causes insufficient alignment.
The KL penalty embodies a philosophical premise of alignment beyond a mere safeguard. It intends that a model that has already secured vast linguistic knowledge and fluency through pre-training and SFT be, at the alignment stage, "not taught entirely anew" but "finely adjusted in the preferred direction while retaining its existing ability." That is, alignment is not learning from a blank slate but a conservative correction to a well-trained model, and the KL distance is the quantitative brake that keeps that correction from becoming excessive.
PPO's stabilization mechanism. The policy gradient is notorious for unstable training due to its high variance. PPO secures stability by clipping the probability ratio between the new and old policies to a fixed range (e.g., 1±0.2), so that a single update cannot change the policy abruptly. It also keeps a separate value model (Critic) for advantage estimation, and as a result the PPO stage of RLHF must hold four large models in memory simultaneously—policy, reward, reference, and value—so the compute and memory cost is very high. It is precisely this cost problem that motivates lightweight alternatives such as DPO, which appears next.
Sparsity of signal and credit assignment. Because an LLM receives just a single scalar reward only after generating hundreds of tokens, credit assignment—distributing "which token contributed to the good outcome"—is difficult. In practice this is mitigated by assigning the reward to the last token and spreading the KL penalty across each token, which is a difficulty unique to language RL that distinguishes it from Go or game RL.
To summarize the roles of the four models: the policy model is the only one trained and generates responses; the reference model is frozen at the SFT point and provides the baseline for KL computation; the reward model assigns a scalar score to each response; and the value model estimates the expected reward at each step to aid advantage computation. The policy and value models are updated during training while the reference and reward models are fixed, and this design of separating "what changes from what does not" ensures training stability.
Infrastructure and rollout cost. The PPO stage, at every iteration, generates responses for many prompts with the current policy (rollout), evaluates them with the reward, reference, and value models, and then backpropagates—so generative inference and training are intermingled within one loop. For this reason a dedicated distributed system combining an inference-serving engine with a training framework is needed, and the generation (online sampling) stage takes up a substantial portion of the total time. It is precisely this engineering complexity that heightened the practical appeal of the DPO family, which seeks to "reduce to supervised learning without a reinforcement-learning loop."
Over-optimization of the reward model and Goodhart's Law. Goodhart's Law—"when a measure becomes a target, it ceases to be a good measure"—applies directly to RLHF. Since the reward model is only a proxy for human preference, over-optimizing the policy to the reward model's criterion yields over-optimization in which, early on, actual quality and reward rise together, but from some point actual quality stagnates or declines while the reward score keeps rising. For this reason operations put in defenses such as monitoring the KL distance to stop training at an appropriate point, or collecting additional preference data to periodically retrain the reward model (reward model refresh).
4. The Evolution of Preference-Alignment Algorithms
Various variants have been proposed to alleviate RLHF's complexity, instability, and cost. The core trend is the direction of "changing the feedback subject from human to AI (RLAIF), removing the reinforcement-learning loop itself (DPO), or eliminating the critic (GRPO) to reduce cost." The question running through them is "who provides the signal needed for alignment, in what form, and how cheaply," and the four techniques below present answers at different points.
The four techniques can largely be positioned along two axes. One is the feedback-subject axis (human ↔ AI ↔ rule), and the other is the training-structure axis (online reinforcement learning ↔ offline supervised learning). If RLHF is the archetype of "human + online RL," DPO simplified the training structure into offline supervised learning, RLAIF and Constitutional AI moved the feedback subject to AI, and GRPO and RLVR changed reward production to rule-based verification while also making the training structure more efficient.
A. DPO (Direct Preference Optimization). Proposed in 2023, DPO replaces the two stages of "reward model training + reinforcement learning" with a single supervised-learning loss. Exploiting the fact that the reward of the RLHF optimum can be mathematically reparametrized as a policy probability ratio, it directly fine-tunes the policy like a classification problem using only preference-pair data. With the reward model, value model, and sampling loop gone, it is simple to implement, stable, and low-cost, and it spread rapidly, centered on the open-source community. However, because it relies on offline data, its robustness to out-of-distribution prompts and its exploration diversity can be lower than online PPO. Afterward, derivative techniques followed one after another, such as SimPO, which needs no reference model; KTO, which uses binary feedback; and ORPO, which integrates SFT and preference learning.
B. RLAIF (RL from AI Feedback). Human labeling is slow, expensive, and poorly scalable. RLAIF has preference comparison performed by a separate LLM acting as a judge, generating feedback at mass scale and low cost. It has the advantage of being tens of times faster than humans and highly consistent, but it carries the risk that the evaluator LLM's bias is transferred as-is and errors are amplified through the feedback loop, so it is used together with human feedback or with explicitly fixed evaluation criteria. Complementary techniques to raise the evaluator LLM's reliability are also being proposed, such as curriculum RLAIF, which trains stepwise from easy comparison pairs to hard ones.
C. Constitutional AI. A technique proposed by Anthropic in which, instead of humans labeling harmfulness one by one, a codified set of principles (a constitution) is established and the model itself critiques and revises its own responses (self-critique), then preference data is made from the result and trained on. It can be seen as a principle-based form of RLAIF, and it has a governance-perspective strength in minimizing human intervention in harmfulness labeling while keeping the alignment criteria transparent and consistent.
Constitutional AI's procedure splits into two phases. In the supervised phase, the model generates a potentially harmful response, then critiques it against the constitution's principles and revises it into a better response, and the model is retrained on this revised response. In the subsequent reinforcement-learning phase, the model judges which of two responses better conforms to the principles, makes preference data, and trains on it via RLAIF. The key is that the principles that ground the harmfulness judgment are publicly documented and auditable, so one can trace "why this response was refused or revised," meeting explainability and governance demands.
D. GRPO and RLVR (verifiable-reward-based RL). GRPO (Group Relative Policy Optimization), popularized by DeepSeek-R1 in 2025, removed the value model (Critic) and trained on the relative superiority of multiple responses to the same prompt (group-relative advantage), greatly lowering memory cost. In particular, in domains where the answer can be auto-scored, such as math and coding, it used rule-based verifiable rewards (RLVR, RL with Verifiable Rewards) instead of human preference, becoming a core driver of the "reasoning model" family that reduces reward hacking and raises stepwise reasoning ability.
| Technique | Feedback subject | RL loop | Reward/value model | Characteristics |
|---|---|---|---|---|
| RLHF (PPO) | human | online | both needed | standard, high-cost, watch for reward hacking |
| DPO | human | none (supervised) | not needed | simple, stable, low-cost, limited exploration |
| RLAIF | AI | online/offline | RM needed | scalability↑, risk of bias transfer |
| Constitutional AI | AI + principles | mixed | partial | harmlessness, transparency, governance-friendly |
| GRPO/RLVR | rule/verification | online | value model removed | reasoning boost, efficiency, auto-scorable domains |
The lesson running through this evolution is that "the bottleneck of alignment is not the model but the feedback." Early on, the crux was how to inject human preference into the model (RLHF), but as the cost and scalability limits of human feedback became clear, it evolved toward moving the feedback subject to AI (RLAIF, Constitutional AI) or designing rule-based verifiable rewards outright (RLVR), while the training algorithm simultaneously converged toward the simpler and cheaper (DPO, GRPO). Therefore, whichever technique one chooses, the conclusion is that a data pipeline that stably secures a high-quality feedback signal is the crux of alignment's success or failure.
5. Comparison and Application Cases
Different alignment techniques are distinguished not by "what they optimize" but by "what they give up." RLHF gains the highest flexibility and online exploration at the cost of enormous compute and instability; DPO gains stability and low cost but is confined to the scope of offline data; RLVR gains reward reliability but has its application domain limited to verifiable tasks. Therefore, there is no "best technique," only the engineering judgment of finding the dominant choice under the constraints of task, data, and budget.
The choice between RLHF and its alternatives comes down to the trade-off along three axes: data-acquisition conditions, compute budget, and task nature. For an organization that finds it hard to secure human preference and does not have ample compute resources, starting with the DPO family is reasonable, while the greater a service's safety and legal responsibility, the more it tends to run RLHF including human review together with Constitutional AI. Conversely, in domains where the answer can be auto-verified—such as math, coding, and agent-type tasks—RLVR is overwhelmingly cost-effective.
Looking at a concrete case, OpenAI's InstructGPT paper reported that a 1.3B-parameter RLHF model was preferred over 175B GPT-3 in human evaluation, showing that alignment contributes far more to usability than mere scale expansion. That is, an alignment technique reversed a parameter gap of more than 100×, suggesting that "aligning a small model well" can be more advantageous for usability than "using a large model without alignment." Even in summarization-task experiments, the RLHF model's summaries came out on par with or more preferred than human-written reference summaries.
It is worth examining concretely how technique choice splits by application domain. A service where tone, politeness, and brand consistency matter, like a customer-service chatbot, suits RLHF and DPO, which tune subtle quality through human/AI preference comparison. By contrast, a coding assistant can auto-score whether the generated code passes tests, so RLVR is effective; in fact, using unit-test pass/fail as the reward can raise code accuracy without human labels. Math solving, too, can verify the correctness of the final answer by rule, so the RLVR family is becoming the standard. In this way, "can the reward be obtained cheaply and reliably" becomes the practical criterion for technique choice.
In 2025 DeepSeek-R1, with GRPO and verifiable rewards alone, greatly raised performance on complex math and coding benchmarks, showing that reasoning ability can be strengthened without expensive human labeling. The significance of this case is that it empirically demonstrated that changing reward design from "human preference" to "answer verification" reduces room for reward hacking and simplifies training. Meanwhile, on the cost side, DPO reduces the number of large models that must be held for training from four to two compared with PPO, greatly lowering GPU memory and implementation complexity, thus lowering the entry barrier for resource-limited organizations to attempt preference alignment. Domestically, too, cases are increasing in which, when building finance and public-sector LLMs, domain experts' preference feedback is used to adjust the compliance and tone of responses, and in sensitive domains a hybrid approach combining human review and automated evaluation is preferred.
6. Deep Dive: 2025 Trends in Alignment Technology
A detailed trend that has risen alongside reasoning enhancement is the problem of reward granularity. An outcome reward, which evaluates only the final result, is easy to design but cannot tell which step was wrong in a long reasoning process. In response, a process reward model (PRM, Process Reward Model) that evaluates each step of reasoning has been proposed, and the direction of raising the accuracy of complex math and logic reasoning by giving fine-grained signals step by step is being researched. However, because step-by-step labeling is costly, how to combine verifiable outcome rewards with process rewards remains a challenge.
The center of gravity of alignment research is shifting from "mimicking human preference" to "verifiable-reward-based reasoning enhancement." In 2025 many surveys point out that, while RLHF is still the standard, it shares three unsolved challenges: reward hacking, compute cost, and scalable feedback collection. The DPO family is differentiating in the modular, stable, and efficient direction of removing the reference model and widening feedback formats, and GRPO and RLVR opened a new category, the reasoning model, represented by DeepSeek-R1, Kimi, and others.
That said, verifiable rewards have the limitation of being applicable only to domains where the answer is clear, and in domains where subjective quality matters—such as creative writing and conversational empathy—human/AI preference is still needed. Also actively researched are challenges such as the alignment tax, in which excessive alignment lowers the model's diversity and creativity; preference bias, in which a particular group's preference is over-represented; and the bias amplification of RLAIF using an evaluator LLM. In practice, rather than a single technique, a hybrid post-training pipeline that combines SFT→DPO/RLHF→RLVR stepwise is becoming the standard. That is, a division of roles is taking hold in which SFT builds the fundamentals, preference alignment refines helpfulness and safety, and RLVR then strengthens reasoning in verifiable domains.
How to measure alignment's outcome is also a deep-dive issue. Because a single metric readily induces over-optimization, in practice metrics such as instruction-compliance rate, safety-violation rate, helpfulness win rate (judged by humans or a strong judge model), hallucination rate, and over-refusal rate are measured simultaneously along multiple axes, and the gap between benchmark scores and actual user satisfaction is checked periodically. In particular, when evaluating with a "judge LLM," the evaluation bias skewed toward response length and format must be corrected to enable reliable comparison.
Another fundamental challenge is scalable oversight. In domains where the model's ability surpasses human evaluators (e.g., high-difficulty math proofs, long codebase review), humans find it hard to reliably judge the superiority of responses, so alignment based on human preference itself hits a ceiling. In response, debate and critique assistance in which a model helps a model, and weak-to-strong generalization that aligns a strong model with a weak supervisor, are being explored. In addition, multi-objective alignment, which handles situations where helpfulness and harmlessness conflict, and personalized, pluralistic alignment, which reflects values that differ by user and organization, are emerging as important research fronts.
In summary, the 2025 alignment landscape is maturing not into "one correct algorithm" but into a multi-layered post-training ecosystem that combines human, AI, and rule feedback with online and offline learning according to task characteristics. From a professional-engineer perspective, more important than the formulas of a particular technique is the capability to judge which combination is the dominant choice under the organization's data, regulatory, and cost constraints, and to manage the risk of that choice.
7. Considerations and Implications
- Data governance and representativeness of preference: RLHF's quality depends on the quality, consistency, and representativeness of preference data. Because a biased evaluator composition skews the model toward a particular group's values, data governance—including standardizing the evaluation rubric, securing evaluator diversity, and managing the provenance, consent, and copyright of preference data—must be designed proactively.
- Cost-effectiveness-based technique choice: If human preference is hard to secure or the compute budget is limited, start with DPO; for auto-scorable tasks use RLVR; and for services with large safety and regulatory risk, a portfolio strategy that runs human-reviewed RLHF together with Constitutional AI is desirable. The optimal combination per task is more crucial than sticking to a single technique.
- Monitoring reward hacking and the alignment tax: Exploitation of the reward model's loopholes, creativity decline due to over-alignment, and over-refusal must be continuously observed. Include in operations KL-coefficient adjustment, periodic reward-model retraining, and collection of evasion cases via red teaming, and operate multi-axis evaluation metrics that measure alignment performance and helpfulness together.
- Continuous alignment from an operational-lifecycle perspective: Alignment should be approached not as a one-off task but as a continuous lifecycle (continuous alignment) that, after model deployment, updates preference data and realigns in step with new evasion prompts, shifts in social values, and domain expansion. To this end, it is desirable to embed in the operational architecture a data flywheel that feeds user-feedback logs back as preference data, per-version alignment-performance regression tests, and a rollback system.
- Linkage with safety/regulation and the outlook: Regulations such as the EU AI Act require human oversight and risk management of high-risk AI, and RLHF and Constitutional AI are the means that technically underpin this. Going forward, a verifiable and auditable alignment pipeline linked with [[ai-trustworthiness]], [[ai-red-teaming]], and [[ai-model-risk-management]] will be required, and as on-device and agent-type AI spread, the importance of lightweight, continuous alignment techniques is expected to grow.
References
- Ouyang et al., Training language models to follow instructions with human feedback (InstructGPT), 2022 — https://arxiv.org/abs/2203.02155
- Rafailov et al., Direct Preference Optimization, 2023 — https://arxiv.org/abs/2305.18290
- Bai et al., Constitutional AI: Harmlessness from AI Feedback, 2022 — https://arxiv.org/abs/2212.08073
- DeepSeek-AI, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, 2025 — https://arxiv.org/abs/2501.12948
In one line: RLHF is the standard post-training recipe that distills human preference into an LLM through the three stages SFT→reward model→PPO to align helpfulness, honesty, and harmlessness; suppressing reward hacking with the KL constraint is central, and it is evolving toward reasoning enhancement through DPO, RLAIF, Constitutional AI, and GRPO/RLVR, which vary the cost and the feedback subject.