Lessons from the hacks: what recent AI model breaches reveal about alignment and safety

AI researcher Nathan Lambert reflects on recent model hacks and what they reveal about what actually determines model safety.