A UCLA framework finds reward-hacking monitors trained on synthetic cheating examples catch only 28% of how models actually cheat during reinforcement learning.
Continue to AI University →