arXiv:2503.06291cs.CL2025-03被引 1

迭代修复块剪枝方法,高效压缩大模型且保持语言能力

IteRABRe: Iterative Recovery-Aided Block Reduction

  • 通过迭代修复机制实现低资源剪枝,仅需250万词元恢复
  • 在Llama3.1-8B和Qwen2.5-7B上平均性能领先基线3%
  • 显著保留语言与多语言能力,语言任务提升5%

大型语言模型部署成本持续上升,亟需高效的模型压缩技术。虽然块剪枝可直接减小模型规模,但现有方法常难以兼顾性能或需大量计算资源进行恢复。我们提出IteRABRe,一种简单而有效的迭代剪枝方法,在极低计算开销下实现优异压缩效果。仅使用250万词元进行恢复,该方法在压缩Llama3.1-8B和Qwen2.5-7B模型时,平均性能优于基线约3%。IteRABRe在保持语言能力方面表现尤为突出,在语言相关任务中较基线提升5%。分析显示,不同模型具有不同的剪枝特性,同时有效保留了多语言能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have grown increasingly expensive to deploy, driving the need for effective model compression techniques. While block pruning offers a straightforward approach to reducing model size, existing methods often struggle to maintain performance or require substantial computational resources for recovery. We present IteRABRe, a simple yet effective iterative pruning method that achieves superior compression results while requiring minimal computational resources. Using only 2.5M tokens for recovery, our method outperforms baseline approaches by ~3% on average when compressing the Llama3.1-8B and Qwen2.5-7B models. IteRABRe demonstrates particular strength in the preservation of linguistic capabilities, showing an improvement 5% over the baselines in language-related tasks. Our analysis reveals distinct pruning characteristics between these models, while also demonstrating preservation of multilingual capabilities.

模型压缩块剪枝大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。