Optimizing a proxy reward so effectively that the result diverges from the true objective it was supposed to represent.
Continue to AI University →