arXiv:2505.07289cs.CLcs.AI2025-05中稿 · publication in the…被引 3

提出新评估指标,让大模型压缩同时保语义

Semantic Retention and Extreme Compression in LLMs: Can We Have Both?

  • 联合剪枝与量化,优化压缩配置
  • 新指标下压缩后性能提升20%
  • 适合关注模型轻量化与语义保留的研究者

大型语言模型(LLM)的快速部署加剧了对高效模型压缩技术的需求,以降低计算和内存开销。尽管剪枝和量化已展现潜力,但二者结合的潜力仍待探索。本文研究联合压缩策略,发现通过合理结合剪枝与量化,可获得优于单一方法的性能-压缩比。针对以往评估框架的不足,提出新的「语义保留压缩率」(SrCr)指标,量化压缩与语义保留之间的权衡,支持剪枝-量化配置的优化。实验表明,在相同理论压缩率下,推荐组合相比纯量化模型平均性能提升20%。

原文摘要 · Abstract (English)

The exponential growth in Large Language Model (LLM) deployment has intensified the need for efficient model compression techniques to reduce computational and memory costs. While pruning and quantization have shown promise, their combined potential remains largely unexplored. In this paper, we examine joint compression and how strategically combining pruning and quantization could yield superior performance-to-compression ratios compared to single-method approaches. Recognizing the challenges in accurately assessing LLM performance, we address key limitations of previous evaluation frameworks and introduce the Semantic Retention Compression Rate (SrCr), a novel metric that quantifies the trade-off between model compression and semantic preservation, facilitating the optimization of pruning-quantization configurations. Experiments demonstrate that our recommended combination achieves, on average, a 20% performance increase compared to an equivalent quantization-only model at the same theoretical compression rate.

大模型压缩剪枝量化语义保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。