Prime Intellect's new Blackwell-native CUDA kernels reportedly speed MoE inference up to 2.4x versus PyTorch's grouped GEMM on B200 GPUs.
Continue to AI University →