UK AI Security Institute finds popular AI safety benchmarks don't measure a consistent trait

UK researchers used psychology-style methods to show popular AI safety benchmarks don't measure one consistent underlying trait.