Apple's DeepAmbigQA benchmark tests whether LLMs find ALL the answers, not just one

Apple researchers built DeepAmbigQA, a benchmark that catches search-equipped LLMs giving incomplete answers to ambiguous multi-hop questions.