Nathan Lambert, author of the RLHF Book, answers technical audience questions on post-training recipes: staged vs. merged training, KL penalties, SFT data scaling, and RL versus in-context learning.
Continue to AI University →