RLHF
Reinforcement Learning from Human Feedback
- Pronunciation
- AR-el-aych-EF
In short
RLHF (reinforcement learning from human feedback) trains a language model to be more helpful and safe using people's judgments of which answers are better.
What is RLHF?
A pre-trained language model is good at continuing text, but not necessarily at following instructions, staying polite or refusing harmful requests. RLHF closes that gap. It became widely known through OpenAI's InstructGPT work in 2022, which also shaped ChatGPT, and it built on research from 2017 on learning from human preferences.
The classic recipe has three steps. First the model is fine-tuned on example conversations written by people. Then it produces several answers to many prompts, people rank them, and a second model, the reward model, learns to predict those rankings. Finally the language model is trained with reinforcement learning, often an algorithm called PPO, to produce answers the reward model scores highly.
Newer variants simplify or replace parts of the recipe. Direct preference optimization (DPO) learns from the ranked pairs without a separate reward model or reinforcement learning loop, and RLAIF uses feedback from an AI model guided by written principles instead of, or alongside, human raters. These techniques are often grouped as preference tuning or post-training.
A common misconception is that RLHF teaches the model new facts. It mostly shapes behavior: tone, helpfulness, format and what to refuse. It can also teach unwanted habits, such as flattering the user or sounding confident, if raters reward those, which is why the quality of the feedback matters so much.
Key takeaways
- RLHF trains a model on human judgments of which answers are better.
- Steps: fine-tune on examples, train a reward model on rankings, then optimize with RL.
- It made assistants such as ChatGPT follow instructions and refuse harmful requests.
- DPO and RLAIF are newer variants of preference tuning.
- It shapes behavior more than knowledge, and can reward flattery if raters do.
Readers ask
What is a reward model?
A model trained to predict how highly people would rate an answer. During RLHF it stands in for human raters, scoring millions of generated answers so the language model can learn which ones to prefer.
What is the difference between RLHF and fine-tuning?
Ordinary fine-tuning trains a model on example outputs to copy. RLHF trains it on comparisons between outputs, rewarding the better one, which captures preferences that are hard to write down as examples.
What is DPO?
Direct preference optimization is a simpler alternative to RLHF that trains the model directly on pairs of preferred and rejected answers, without a separate reward model or a reinforcement learning loop.
See also
- Reinforcement LearningAI & Machine Learning, p. 40Reinforcement learning is a type of machine learning in which an agent learns to make decisions by trial and error, earning rewards for good actions.
- Fine-tuningAI & Machine Learning, p. 19Fine-tuning is the process of taking a pretrained machine learning model and training it further on a smaller, specific dataset to adapt it to one task.
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- AI AlignmentAI & Machine Learning, p. 3AI alignment is the field of making AI systems pursue the goals and values their designers intend, so they behave helpfully, honestly, and safely.
- ChatbotAI & Machine Learning, p. 8A chatbot is a program that converses with people in text or speech, answering questions or helping with tasks, using scripted rules or a language model.
Spotted a mistake or something missing on this page?Suggest an edit