New AA-AnalystAgent Benchmark: Claude Opus 5 Beats GPT-5.5 on Spreadsheet Agent Reliability

Artificial Analysis's new AA-AnalystAgent benchmark tests AI agents on real spreadsheet tasks — Claude Opus 5 leads, but even the best model only nails every task 54% of the time.