This was a data center a year ago… Now it's on my desk

A short hands-on demo of local inference throughput on an ASUS workstation, scaling toward roughly 2,600 tokens per second as concurrent requests increase.