arXiv:2601.02674cs.CL2026-01Conference of the …被引 3

通过多领域校准迭代剪枝,大幅压缩大模型体积且保持性能

Iterative Structured Pruning for Large Language Models with Multi-Domain Calibration

  • 采用多领域校准数据集与迭代策略识别冗余通道
  • 在多个下游任务上实现显著压缩,性能损失极小
  • 适合需要高效部署大模型的工程应用

大型语言模型(LLMs)在自然语言处理任务中取得了显著成果,但其不断增长的规模带来了计算开销、内存占用和推理延迟等实际部署障碍。模型剪枝是应对这些挑战的有效方案,但现有无结构剪枝常产生不规则稀疏模式,需专用硬件或软件支持。本文探索结构化剪枝,通过移除完整的网络组件来保持与标准硬件加速器的兼容性。提出一种新型结构化剪枝框架,结合混合多领域校准集与迭代校准策略,有效识别并移除冗余通道。在多种模型和下游任务上的大量实验表明,该方法实现了显著压缩,同时性能损失极小。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable success across a wide spectrum of natural language processing tasks. However, their ever-growing scale introduces significant barriers to real-world deployment, including substantial computational overhead, memory footprint, and inference latency. While model pruning presents a viable solution to these challenges, existing unstructured pruning techniques often yield irregular sparsity patterns that necessitate specialized hardware or software support. In this work, we explore structured pruning, which eliminates entire architectural components and maintains compatibility with standard hardware accelerators. We introduce a novel structured pruning framework that leverages a hybrid multi-domain calibration set and an iterative calibration strategy to effectively identify and remove redundant channels. Extensive experiments on various models across diverse downstream tasks show that our approach achieves significant compression with minimal performance degradation.

大模型剪枝结构化剪枝模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。