Model alignment · 2022 · Long Ouyang et al.
Training Language Models to Follow Instructions with Human Feedback
Separate pretraining capability from assistant behavior using demonstrations, a learned preference model, and constrained reinforcement learning.
The central move
Separate pretraining capability from assistant behavior using demonstrations, a learned preference model, and constrained reinforcement learning.
Why it had to exist
Next-token prediction produces broad capability but does not reliably optimize helpful instruction following, truthfulness, or harmless behavior for a user-facing assistant.
Where it leads
Pretraining → demonstrations and preference learning → assistant post-training, safety evaluation, and governance.