让大模型压缩更懂损失,提升压缩后性能。
LaMoC: Loss-Aware Modular Compression for LLMs

- 融合激活统计与经验费舍尔信息,动态选择压缩特征
- 在4-8B模型上降低2.5%困惑度,任务准确率提升1%
- 适合追求高效压缩且不牺牲性能的模型部署者
模块化压缩已显著减少大语言模型参数量,同时保持强大的语言理解能力与下游任务精度。然而,现有联合模块压缩方法主要依赖激活统计,对损失敏感性信息及其模块级表征未充分探索。本文提出LaMoC,一种损失感知的模块化压缩方法,通过梯度误差对齐融合激活统计与经验费舍尔信息。该方法改进联合压缩,选择能更好对齐局部模块重建误差与下游损失的压缩统计量。贡献有三:(1) 将经验费舍尔作为模块级损失感知代理,可与压缩所需的激活统计融合;(2) 将联合模块压缩重构为双层优化问题,最小化模块重建误差,同时调节激活与梯度信息融合率;(3) 实现基于统计验证的实证驱动方法求解压缩问题。我们在四种模型家族共八种模型上评估了LaMoC。在4-8B模型上,相比最先进方法,平均困惑度降低2.5%,任务准确率相对提升1%。
原文摘要 · Abstract (English)
Modular compression has enabled considerable parameter reduction in LLMs while preserving strong language understanding and downstream task accuracy. However, existing joint modular compression methods primarily rely on activation statistics, leaving loss-sensitivity information and its module-level characterization underexplored. We investigate addressing this gap with LaMoC, a loss-aware modular compression methodology that blends activation and Empirical Fisher statistics through gradient-error alignment. LaMoC improves joint compression by selecting compression statistics that better align local module reconstruction error with the downstream loss. Our contributions are three-fold: (1) We characterize the Empirical Fisher as a module-level loss-aware proxy that can be blended with the activation statistics required for compression. (2) We reformulate joint modular compression as a two-tiered optimization problem that minimizes module reconstruction error while tuning the activation and gradient information blending rate. (3) We implement an empirically driven methodology with statistical validation to solve the resulting compression problem. We evaluate LaMoC across four model families spanning eight models. On the 4-8B models, LaMoC achieves an average 2.5% reduction in perplexity and a 1% relative improvement in task accuracy over state-of-the-art modular compression methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。