Over-Optimization and RLHF’s Bad Reputation | Post-Training Course, Lecture 9
Nathan Lambert explains reward-model over-optimization using OpenAI's own GPT-4o sycophancy postmortem as the case study, then traces how RLHF's early reputation for being 'just style transfer' gave…