The proportion of a training corpus allocated to each data domain (code, web text, math, dialogue, etc.), tuned as a deliberate recipe rather than left to raw source volume.
Continue to AI University →