A new attention method runs up to 11x faster than FlashAttention-2 on long context while barely losing recall accuracy, using a quarter of the memory.
Continue to AI University →