Accelerating Text-to-Video Generation With Calibrated Sparse Attention

Apple researchers speed up text-to-video diffusion models by skipping attention computations that barely affect the output.