OpenAI's self-play agent GPT-Red found more prompt-injection exploits than human red-teamers, and trained GPT-5.6 against them.
Continue to AI University →