Apple researchers built DeepAmbigQA, a benchmark that catches search-equipped LLMs giving incomplete answers to ambiguous multi-hop questions.
Continue to AI University →