Researchers Say They Extracted a Frontier Model's Hidden Reasoning by Replaying It Into a Weaker Model
A new paper claims researchers replayed a frontier model's encrypted chain-of-thought into a weaker sibling model, jailbroke that model, and recovered the stronger model's hidden reasoning in plainte…