Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity
arXiv:2512.12744v4 Announce Type: replace Abstract: Activation sparsity offers a compelling route to accelerate large language model (LLM) inference by selectively suppressing hidden activations, yet existing approaches exhibit severe accuracy degradation at high sparsity. We show that this failure stems from…
