Skip to main content

RLHF

Reinforcement Learning from Human Feedback

Pronunciation
AR-el-aych-EF
Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/rlhf

In short

RLHF (reinforcement learning from human feedback) trains a language model to be more helpful and safe using people's judgments of which answers are better.

What is RLHF?

A pre-trained language model is good at continuing text, but not necessarily at following instructions, staying polite or refusing harmful requests. RLHF closes that gap. It became widely known through OpenAI's InstructGPT work in 2022, which also shaped ChatGPT, and it built on research from 2017 on learning from human preferences.

The classic recipe has three steps. First the model is fine-tuned on example conversations written by people. Then it produces several answers to many prompts, people rank them, and a second model, the reward model, learns to predict those rankings. Finally the language model is trained with reinforcement learning, often an algorithm called PPO, to produce answers the reward model scores highly.

Newer variants simplify or replace parts of the recipe. Direct preference optimization (DPO) learns from the ranked pairs without a separate reward model or reinforcement learning loop, and RLAIF uses feedback from an AI model guided by written principles instead of, or alongside, human raters. These techniques are often grouped as preference tuning or post-training.

A common misconception is that RLHF teaches the model new facts. It mostly shapes behavior: tone, helpfulness, format and what to refuse. It can also teach unwanted habits, such as flattering the user or sounding confident, if raters reward those, which is why the quality of the feedback matters so much.

Key takeaways

  • RLHF trains a model on human judgments of which answers are better.
  • Steps: fine-tune on examples, train a reward model on rankings, then optimize with RL.
  • It made assistants such as ChatGPT follow instructions and refuse harmful requests.
  • DPO and RLAIF are newer variants of preference tuning.
  • It shapes behavior more than knowledge, and can reward flattery if raters do.

Readers ask

What is a reward model?

A model trained to predict how highly people would rate an answer. During RLHF it stands in for human raters, scoring millions of generated answers so the language model can learn which ones to prefer.

What is the difference between RLHF and fine-tuning?

Ordinary fine-tuning trains a model on example outputs to copy. RLHF trains it on comparisons between outputs, rewarding the better one, which captures preferences that are hard to write down as examples.

What is DPO?

Direct preference optimization is a simpler alternative to RLHF that trains the model directly on pairs of preferred and rejected answers, without a separate reward model or a reinforcement learning loop.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings