OpenAI Finds Reliability Problems in the SWE-Bench Pro Coding Benchmark

OpenAI's own analysis shows a popular coding benchmark, SWE-Bench Pro, has flaws that can misjudge how good AI models really are at coding.