SGLang and Meta Speed Up LLM Prefill Up to 1.93x With a Redesigned CUDA Graph

SGLang and Meta rebuilt CUDA Graph support to cut LLM prefill time by up to 1.93x with faster startup and broader kernel support.