Skip to content

What is…

RLHF (Reinforcement Learning from Human Feedback)

Training a model on human ratings of its answers so it becomes more helpful, honest and safe.

After pre-training, people compare pairs of model answers and pick the better one. Those preferences train a "reward model," which then steers the main model toward answers humans prefer. It's a big part of why assistants are polite, follow instructions and decline harmful requests.

Variants use AI feedback guided by a written set of principles (Anthropic's "Constitutional AI") or skip the reward model entirely (DPO).

💡 Think of it like

A comedian refining their set based on which jokes got laughs.

🧠 Test yourself

Which of these describes RLHF?

Related terms

🎮 Learn AI by playing

40 bite-size missions, boss battles and a certificate. Free.

Start the bootcamp →

☀️ One AI term every morning

Plus the day's top stories, in your inbox by 8am.