Models and training

Reinforcement learning from human feedback (RLHF)

stable definition

Definition

A model-alignment method that uses human preference data to train a reward signal or otherwise optimize model behavior toward preferred responses.

Distinguish it from nearby terms

RLHF shapes model behavior during training; human-in-the-loop describes human participation in an operating workflow.

Check your understanding

Preference optimization does not guarantee truthfulness, policy compliance, or safe tool use.

Also called

RLHF

Public evidence