Meta's new WildArtifactBench judges AI agents on messy, real-world tasks using human and AI preference votes instead of a fixed scoring rubric.
Continue to AI University →