Reinforcement learning from human feedback
Training a model against human judgements of which answer is better, rather than against a single correct answer.
Also written: RLHF, preference tuning, alignment training.
For most of what an assistant does there is no single right answer, so there is nothing to train against directly. Instead people compare pairs of answers, a second model learns to predict those preferences, and the first model is trained to score well on it.
This is where tone, helpfulness and refusals come from. It is also where some of the characteristic failures come from: a model trained to produce answers people prefer learns, among other things, that people prefer confident answers.