提出新稀疏性机制与加速系统,显著提升大模型参数高效微调速度。
Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
- 识别并建模新型稀疏性‘Shadowy Sparsity’,增强微调时的稀疏特征捕捉。
- 在端到端微调中实现最高2.49倍加速,显著降低时间与成本。
- 适合关注大模型高效微调、硬件优化的研究者与工程团队。
通过微调将预训练大语言模型(LLMs)适配至多样化下游任务至关重要,但参数高效微调(PEFT)技术效率低下,带来显著的时间投入和运营成本。本文首次引入一种新型稀疏性——‘Shadowy Sparsity’,其在微调过程中具有独特性且未被充分研究。针对该稀疏性,我们提出 Long Exposure 系统以加速 LLM 的 PEFT。该系统包含三个核心组件:Shadowy-sparsity Exposer 采用延长感知范围,更精准捕获阴影稀疏下的细节;Sequence-oriented Predictor 实现对长序列输入与持续演化的参数的高效准确预测;Dynamic-aware Operator 则通过结构化计算模式与合并内存访问,优化动态稀疏操作。大量实验表明,Long Exposure 在端到端微调中相较现有方法最高实现2.49倍加速,为 LLM 的 PEFT 加速提供了极具前景的解决方案。
原文摘要 · Abstract (English)
The adaptation of pre-trained large language models (LLMs) to diverse downstream tasks via fine-tuning is critical for numerous applications. However, the inefficiency of parameter-efficient fine-tuning (PEFT) techniques presents significant challenges in terms of time investments and operational costs. In this paper, we first introduce a nuanced form of sparsity, termed Shadowy Sparsity, which is distinctive in fine-tuning and has not been adequately addressed for acceleration. Under Shadowy Sparsity, we propose Long Exposure, an efficient system to accelerate PEFT for LLMs. Long Exposure comprises three key components: Shadowy-sparsity Exposer employs a prolonged sensing range to capture more sparsity details under shadowy sparsity; Sequence-oriented Predictor provides efficient yet accurate predictions to handle large sequence inputs and constantly-evolving parameters; and Dynamic-aware Operator facilitates more structured computational patterns and coalesced memory accesses, addressing dynamic sparse operations. Extensive evaluations show that Long Exposure outperforms state-of-the-arts with up to a $2.49\times$ speedup in end-to-end fine-tuning, offering promising advancements in accelerating PEFT for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。