SPDP Combines Static and Dynamic Weight Pruning for Faster LLM Inference on GPUs

A new GPU inference framework combines permanent weight pruning with input-adaptive pruning to speed up LLM serving without the usual runtime overhead.