A training trick where the model is fed the correct previous token instead of its own last guess, so early mistakes can't derail the rest of the sequence.
Continue to AI University →