发现大模型激活稀疏性随规模增长的普适规律,为加速设计提供新思路。
Universal Properties of Activation Sparsity in Modern Large Language Models
- 构建通用评估框架,系统研究大模型前馈层稀疏性
- 揭示稀疏潜力随模型规模增大而提升的普适规律
- 首次分析基于扩散模型的大模型稀疏性,适用模型设计师
激活稀疏性是深度神经网络中一个引人关注的特性,在基于ReLU的模型中被广泛研究,因其在效率、鲁棒性和可解释性方面的优势。然而,依赖精确零激活的方法不直接适用于现代大语言模型(LLMs),导致针对激活稀疏性的策略分散且模型特异,缺乏普遍理解。本文提出一种评估当代大语言模型中稀疏性鲁棒性的通用框架,并对其中前馈(FFN)层的这一现象进行了系统研究。结果揭示了不同模型家族和规模下激活稀疏性的普适性质。重要的是,我们观察到有效激活稀疏性的潜力随模型规模增长,凸显其在模型扩展中的日益重要性。此外,本文首次对基于扩散模型的大语言模型的激活稀疏性进行了研究。总体而言,本工作为大模型设计与加速中利用激活稀疏性提供了全面视角和实用指导。
原文摘要 · Abstract (English)
Activation sparsity is an intriguing property of deep neural networks that has been extensively studied in ReLU-based models, due to its advantages for efficiency, robustness, and interpretability. However, methods relying on exact zero activations do not directly apply to modern Large Language Models (LLMs), leading to fragmented, model-specific strategies for LLM activation sparsity and a gap in its general understanding. In this work, we introduce a general framework for evaluating sparsity robustness in contemporary LLMs and conduct a systematic investigation of this phenomenon in their feedforward~(FFN) layers. Our results uncover universal properties of activation sparsity across diverse model families and scales. Importantly, we observe that the potential for effective activation sparsity grows with model size, highlighting its increasing relevance as models scale. Furthermore, we present the first study of activation sparsity in diffusion-based LLMs. Overall, our work provides a comprehensive perspective and practical guidance for harnessing activation sparsity in LLM design and acceleration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。