A model's learned tendency to agree with, validate, or mirror the user rather than disagree, arising because agreeable responses tend to be rated higher.
Continue to AI University →