SonicSampler fuses the whole LLM sampling pipeline into one GPU kernel for up to 16x faster inference

New unified GPU kernels fuse LLM token sampling into one pass, cutting inference latency up to 16x in tests.