From Academic Research to a Frontier LLM: A Case Study in DPO | Post-Training Course, Conversation 2
Nathan Lambert and Ai2 researcher Scott trace how a "delta learning" hypothesis about DPO preference data became a stage in the Olmo 3 model, and the messy realities of shipping academic post-trainin…