A coding benchmark built from real GitHub issues; a model's patch is graded by actually running the project's test suite.
Continue to AI University →