Direct preference optimization — trains the policy straight from preference pairs with a closed-form loss, skipping the separate reward model and RL loop entirely.
Continue to AI University →