Netflix Details Its Internal LLM Serving Platform Built on Triton and vLLM

Netflix shared production lessons from building an in-house LLM inference platform, covering how it handles varied model sizes, hardware, and a fast-moving inference-engine landscape.