Skip to content

What is…

Alignment

Making AI systems reliably do what people actually intend — helpful, honest, and not harmful.

Alignment is the research field (and the practical work) of getting models to pursue the goals we mean, not just the goals we literally typed — and to refuse genuinely harmful requests. RLHF, constitutional training, red-teaming and interpretability research all fall under it.

It matters more as models gain autonomy: an agent with access to your email should be very sure what you want before it acts.

🧠 Test yourself

Which of these describes Alignment?

Related terms

🎮 Learn AI by playing

40 bite-size missions, boss battles and a certificate. Free.

Start the bootcamp →

☀️ One AI term every morning

Plus the day's top stories, in your inbox by 8am.