SGLang and Meta rebuilt CUDA Graph support to cut LLM prefill time by up to 1.93x with faster startup and broader kernel support.
Continue to AI University →