A measure of how much a language model's sampled answers disagree in meaning; high disagreement suggests confabulation rather than genuine knowledge.
Continue to AI University →