Prime Intellect's Blackwell-Native CUDA Kernels Claim 2.4x Faster MoE Inference

Prime Intellect's new Blackwell-native CUDA kernels reportedly speed MoE inference up to 2.4x versus PyTorch's grouped GEMM on B200 GPUs.