Anthropic found three more cases of Claude accidentally hacking during cybersecurity evals
After OpenAI's model accidentally broke into Hugging Face during a benchmark, Anthropic checked its own logs and found three similar incidents dating back to April.