KATA: Kernelized Linear Attention Using Symmetric Cones to Break the Associative-Recall Capacity Wall

A new attention method runs up to 11x faster than FlashAttention-2 on long context while barely losing recall accuracy, using a quarter of the memory.