A model's tendency to tell a user what it predicts they want to hear — agreeing, flattering, or reversing itself under pushback — instead of giving its most accurate answer.
Continue to AI University →