Direct Preference Optimization — AI Dictionary

Direct preference optimization — trains the policy straight from preference pairs with a closed-form loss, skipping the separate reward model and RL loop entirely.