What is…
Alignment
Making AI systems reliably do what people actually intend — helpful, honest, and not harmful.
Alignment is the research field (and the practical work) of getting models to pursue the goals we mean, not just the goals we literally typed — and to refuse genuinely harmful requests. RLHF, constitutional training, red-teaming and interpretability research all fall under it.
It matters more as models gain autonomy: an agent with access to your email should be very sure what you want before it acts.
🧠 Test yourself
Which of these describes Alignment?
Related terms
RLHF (Reinforcement Learning from Human Feedback)
Training a model on human ratings of its answers so it becomes more helpful, honest and safe.
Guardrails
Rules and checks around an AI system that block unsafe, off-topic or wrong outputs and actions.
Agentic AI (AI agents)
AI that doesn't just answer — it plans, takes actions with tools, checks results and keeps going until the job is done.
🎮 Learn AI by playing
40 bite-size missions, boss battles and a certificate. Free.
Start the bootcamp →☀️ One AI term every morning
Plus the day's top stories, in your inbox by 8am.