UK researchers used psychology-style methods to show popular AI safety benchmarks don't measure one consistent underlying trait.
Continue to AI University →