Support Basis: Fast Attention Beyond Bounded Entries
arXiv:2510.01643v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks. However, the quadratic complexity of softmax attention remains a central bottleneck that limits their scalability. Alman and Song (NeurIPS 2023a; NeurIPS…
