SpecPrefetch predicts which MoE experts you'll need next so inference doesn't stall waiting for them

The paper reports an adapter that prefetches mixture-of-experts weights before routing, boosting on-device decoding throughput up to 20% on a phone-class chip.