New vLLM Plugin Splits Attention From Expert Compute to Cut DeepSeek-V3.2 Latency 47%

A new vLLM plugin separates attention from expert computation, cutting DeepSeek-V3.2 response times by 47%.