vLLM's new adaptive speculative-decoding tuner beats every fixed config on DeepSeek-V4, holding peak throughput from 1 to 256 concurrent requests.
Continue to AI University →