QCon talk: how to cut LLM inference costs by an order of magnitude for high-volume workloads

A QCon San Francisco talk lays out hardware, runtime, and queueing trade-offs for building the cheapest possible LLM inference stack.