Using a language model to score or compare other models' outputs instead of relying on human raters.
Continue to AI University →