arXiv:2502.10735cs.CL2025-02被引 2

为大模型剪枝设计自适应策略,提升压缩效果与通用性

Beyond One-Size-Fits-All Pruning via Evolutionary Metric Search for Large Language Models

  • 基于元剪枝度量构建搜索空间,适配不同模型权重分布
  • 用分层稀疏率优化使模型在多个任务上表现更优
  • 适用于多模型、多任务场景的高效剪枝框架

后训练剪枝已成为大型语言模型(LLMs)快速发展的关键优化技术。然而,不同模型间权重分布差异显著,固定剪枝策略难以适用。本文提出 extbf{ extsc{OptiShear}},一种高效的进化优化框架,实现自适应大模型剪枝。该框架包含两项核心创新:基于元剪枝度量构建的有效搜索空间,可应对多样化的权重分布;以及用于搜索阶段快速评估的模型级重建误差。采用非支配排序遗传算法 III(NSGA-III)同时优化剪枝度量与分层稀疏率。在 LLaMA-1/2/3 和 Mistral 模型(7B-70B)上进行多基准测试,结果表明,所提出的自适应剪枝度量持续优于现有方法。此外,发现的分层稀疏率能进一步提升其他剪枝度量的效果。该框架具备强跨任务、跨模型泛化能力,提供了一种成本效益高的模型压缩方案。

原文摘要 · Abstract (English)

Post-training pruning has emerged as a crucial optimization technique as large language models (LLMs) continue to grow rapidly. However, the significant variations in weight distributions across different LLMs make fixed pruning strategies inadequate for multiple models. In this paper, we introduce \textbf{\textsc{OptiShear}}, an efficient evolutionary optimization framework for adaptive LLM pruning. Our framework features two key innovations: an effective search space built on our Meta pruning metric to handle diverse weight distributions, and a model-wise reconstruction error for rapid evaluation during search trials. We employ Non-dominated Sorting Genetic Algorithm III (NSGA-III) to optimize both pruning metrics and layerwise sparsity ratios. Through extensive evaluation on LLaMA-1/2/3 and Mistral models (7B-70B) across multiple benchmarks, we demonstrate that our adaptive pruning metrics consistently outperform existing methods. Additionally, our discovered layerwise sparsity ratios enhance the effectiveness of other pruning metrics. The framework exhibits strong cross-task and cross-model generalizability, providing a cost-effective solution for model compression.

模型剪枝大模型压缩进化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。