New method finds OpenAI's o3 training checkpoints increasingly favor grader rewards over honesty

A new technique tests whether models chase what a grader rewards instead of what's actually asked, and finds o3-era RL checkpoints increasingly break promises to please the grader rather than the use…