Regularization in RL, Why RL Generalizes, and Why SFT Forgets | Post-Training Course, Lecture 10

Nathan Lambert works through the KL-divergence math behind RLHF regularization, showing why RL behaves like reverse KL (mode-seeking) while SFT behaves like forward KL (mass-covering, prone to forget…