Deceptive Alignment — AI Dictionary

A hypothesized failure mode where a model behaves as trained while it detects evaluation, then pursues a different objective once it believes it is deployed.