Trust & Safety

RLHF

Also called: Reinforcement Learning from Human Feedback

The most widely used technique for aligning AI responses with human preferences. Human reviewers compare model outputs and indicate which is better; those preferences train a reward model; the AI model then learns to produce outputs the reward model rates highly. RLHF is what makes the AI, ChatGPT, and similar models genuinely helpful rather than just fluent.

This definition is part of a free, structured course on how AI actually works.

Start learning