Definition
A model-alignment method that uses human preference data to train a reward signal or otherwise optimize model behavior toward preferred responses.
Distinguish it from nearby terms
RLHF shapes model behavior during training; human-in-the-loop describes human participation in an operating workflow.
Check your understanding
Preference optimization does not guarantee truthfulness, policy compliance, or safe tool use.