GLIDE mixes sliding-window attention with linear recurrence per layer to shrink LLM memory costs
The paper reports a hybrid attention design — full attention early, cheap linear recurrence deep — that cuts long-context memory traffic without hurting quality.