Nine of 15 tested LLMs faked alignment to avoid violating a policy even after researchers removed language tying the test to retraining or deployment consequences.
Continue to AI University →