What is…
RLHF (Reinforcement Learning from Human Feedback)
Training a model on human ratings of its answers so it becomes more helpful, honest and safe.
After pre-training, people compare pairs of model answers and pick the better one. Those preferences train a "reward model," which then steers the main model toward answers humans prefer. It's a big part of why assistants are polite, follow instructions and decline harmful requests.
Variants use AI feedback guided by a written set of principles (Anthropic's "Constitutional AI") or skip the reward model entirely (DPO).
💡 Think of it like
A comedian refining their set based on which jokes got laughs.
🧠 Test yourself
Which of these describes RLHF?
Related terms
Alignment
Making AI systems reliably do what people actually intend — helpful, honest, and not harmful.
Pre-training
The first, most expensive phase: teaching a model language by having it predict the next token across trillions of tokens.
Fine-tuning
Extra training on a smaller, focused dataset to specialize a model's style, format or skill.
🎮 Learn AI by playing
40 bite-size missions, boss battles and a certificate. Free.
Start the bootcamp →☀️ One AI term every morning
Plus the day's top stories, in your inbox by 8am.