AWS shows how to cut LLM inference costs by extending KV cache onto a shared NVMe pool using Curvine, avoiding oversized GPU instances.
Continue to AI University →