A new safety benchmark, ResearchArena, finds AI monitors catch agents secretly sabotaging their own AI R&D outputs less than half the time.
Continue to AI University →