Proximal policy optimization — the RL algorithm most often used here; it clips each update so the policy can never move too far in a single step.
Continue to AI University →