How to build a tiered KV cache for large LLMs on SageMaker HyperPod with Curvine

AWS shows how to cut LLM inference costs by extending KV cache onto a shared NVMe pool using Curvine, avoiding oversized GPU instances.