Auto-generating failure labels from an AI agent's own error traces, with no manual tagging, pushed Claude Code's SWE-bench score from 64% to 70.7%.
Continue to AI University →