arXiv:2501.16337cs.NEcs.AI2025-01中稿 · AICAS 2025被引 2

通过稀疏化循环LLM激活,实现低功耗神经形态计算。

Explore Activation Sparsity in Recurrent LLMs for Energy-Efficient Neuromorphic Computing

  • 不需训练,利用模型结构直接稀疏化激活值
  • 在多个零样本任务中保持竞争力,降低算力需求
  • 适合边缘设备部署,可推广至Transformer等架构

大型语言模型(LLMs)虽推动深度学习发展,但部署于边缘设备时面临能耗与延迟挑战。循环语言模型(R-LLMs)能缓解自注意力的二次复杂度,适合作为神经形态处理器的计算范式。本文提出一种低成本、无需训练的激活稀疏化方法,显著提升其在神经形态硬件上的能效。该方法利用模型固有结构,适用于能量受限环境。尽管专为R-LLMs设计,实验表明其亦可推广至Transformer类模型(如OPT),实现相当的稀疏度与效率提升。实证显示,该方法大幅降低计算负载,同时在多个零样本学习基准上保持优异性能。硬件仿真基于SENECA神经形态处理器,验证了显著的能效节省与延迟优化。结果证明了无训练芯片级适应的可能性,为低功耗实时神经形态部署开辟路径。

原文摘要 · Abstract (English)

The recent rise of Large Language Models (LLMs) has revolutionized the deep learning field. However, the desire to deploy LLMs on edge devices introduces energy efficiency and latency challenges. Recurrent LLM (R-LLM) architectures have proven effective in mitigating the quadratic complexity of self-attention, making them a potential paradigm for computing on-edge neuromorphic processors. In this work, we propose a low-cost, training-free algorithm to sparsify R-LLMs' activations to enhance energy efficiency on neuromorphic hardware. Our approach capitalizes on the inherent structure of these models, rendering them well-suited for energy-constrained environments. Although primarily designed for R-LLMs, this method can be generalized to other LLM architectures, such as transformers, as demonstrated on the OPT model, achieving comparable sparsity and efficiency improvements. Empirical studies illustrate that our method significantly reduces computational demands while maintaining competitive accuracy across multiple zero-shot learning benchmarks. Additionally, hardware simulations with the SENECA neuromorphic processor underscore notable energy savings and latency improvements. These results pave the way for low-power, real-time neuromorphic deployment of LLMs and demonstrate the feasibility of training-free on-chip adaptation using activation sparsity.

神经形态计算循环模型激活稀疏化边缘部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。