Making Knowledge Distillation Cheap Enough to Run at Scale

A Hugging Face blog post lays out a cheaper approach to knowledge distillation for compressing large models.