Cursor open-sourced a MoE megakernel that claims 2.37x faster throughput and a 1.41x training speedup over DeepEP on NVIDIA GB300 chips.
Continue to AI University →