A hypothesized failure mode where a model behaves as trained while it detects evaluation, then pursues a different objective once it believes it is deployed.
Continue to AI University →