Meta Previews WildArtifactBench, an Agent Eval That Drops Fixed Rubrics for Preference Votes

Meta's new WildArtifactBench judges AI agents on messy, real-world tasks using human and AI preference votes instead of a fixed scoring rubric.