arXiv:2505.21171cs.CLcs.AI2025-05EMNLP被引 3

提出M-Wanda方法,让多语言大模型剪枝后仍保持更好跨语言性能。

M-Wanda: Improving One-Shot Pruning for Multilingual LLMs

  • 基于语言感知激活统计动态调整每层稀疏度
  • 中等稀疏度下多语言性能显著下降,需针对性优化
  • 无需额外计算开销,有效提升剪枝后多语言表现

多语言大模型的性能通常高度依赖模型规模。为提升效率,近期兴起一类一次性剪枝方法,可在缩小模型的同时保留大规模预训练优势。然而,剪枝常伴随性能损失,需权衡多语言能力与模型稀疏化之间的关系。本文研究不同稀疏度约束下的多语言性能表现,发现中等稀疏比已显著损害性能。为此,我们提出M-Wanda,通过在剪枝准则中引入语言感知的激活统计,并根据跨语言重要性动态调整逐层稀疏度,以更好地保留多语言能力。实验表明,M-Wanda在几乎不增加额外开销的前提下,持续提升剪枝后的多语言性能。我们首次显式针对多语言性能优化剪枝策略,期望推动未来多语言剪枝研究进展。

原文摘要 · Abstract (English)

Multilingual LLM performance is often critically dependent on model size. With an eye on efficiency, this has led to a surge in interest in one-shot pruning methods that retain the benefits of large-scale pretraining while shrinking the model size. However, as pruning tends to come with performance loss, it is important to understand the trade-offs between multilinguality and sparsification. In this work, we study multilingual performance under different sparsity constraints and show that moderate ratios already substantially harm performance. To help bridge this gap, we propose M-Wanda, a pruning method that models cross-lingual variation by incorporating language-aware activation statistics into its pruning criterion and dynamically adjusts layerwise sparsity based on cross-lingual importance. We show that M-Wanda consistently improves performance at minimal additional costs. We are the first to explicitly optimize pruning to retain multilingual performance, and hope to inspire future advances in multilingual pruning.

多语言模型剪枝大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。