SFT — AI Dictionary

Supervised fine-tuning — training the policy to imitate human-written demonstrations, before any reward model is involved.