Google DeepMind: Debate Training Reduces Reward Hacking When Training Against an LLM Judge

A GDM Alignment paper finds that adding a debate opponent during RL training cuts down on LLM judges getting fooled into over-rewarding.