New Benchmark Finds Frontier AI Agents Get Pathogen-Surveillance Analysis Right Only About Half the Time

BioSecBench-Surveillance tested 16 model-harness combos on real genomic-surveillance analysis tasks — the best, Opus 4.8 and GPT-5.5 with Codex, both topped out around 50% accuracy.