The latency spike while a newly started replica loads model weights into GPU memory, before it can serve a request.
Continue to AI University →