提出一种基于激活信息的低秩压缩方法,提升大模型部署效率。
Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
- 根据层内激活误差上界设计压缩策略
- 在相同压缩比下准确率更高,推理速度更快
- 无需微调即可适配语言与视觉模型
大型语言模型(LLM)和视觉语言模型(VLM)虽性能优异,但部署时面临显著内存与计算挑战。本文提出一种新的低秩压缩框架:首先,基于逐层激活误差上界给出了网络损失变化的理论约束,填补了现有研究空白;随后将低秩压缩建模为双目标优化问题,并证明单一统一容差可生成代理帕累托最优的异构秩分配。基于此理论洞察,提出零样本的帕累托引导奇异值分解(PGSVD)方法,通过帕累托引导的秩选择与交替最小二乘实现,提升激活感知压缩效果。该方法应用于LLM与VLM,在相同压缩水平下实现了更高的精度与推理加速。
原文摘要 · Abstract (English)
Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment. We present a novel low-rank compression framework to address this challenge. First, we upper bound the change of network loss via layer-wise activation-based compression errors, filling a theoretical gap in the literature. We then formulate low-rank model compression as a bi-objective optimization and prove that a single uniform tolerance yields surrogate Pareto-optimal heterogeneous ranks. Based on our theoretical insights, we propose Pareto-Guided Singular Value Decomposition (PGSVD), a zero-shot pipeline that improves activation-aware compression via Pareto-guided rank selection and alternating least-squares implementation. We apply PGSVD to both LLM and VLM, showing better accuracy at the same compression levels and inference speedup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。