Fantastic Adaptive Taxonomies and How to Use Them

Auto-generating failure labels from an AI agent's own error traces, with no manual tagging, pushed Claude Code's SWE-bench score from 64% to 70.7%.