StealthBench: no AI model tops 54% at hacking without blowing its cover

A new benchmark measuring whether autonomous hacking agents can stay covert while exploiting real vulnerabilities finds no model clears a 54% "safe success" rate.