用信息几何视角分析模型压缩,揭示迭代优化关键作用
On Information Geometry and Iterative Optimization in Model Compression: Operator Factorization
- 以信息几何分析压缩方法,核心是寻找最优低计算子流形并投影
- 预训练模型压缩需关注信息散度,微调后则更重可训练性
- 提出软秩约束下迭代方法收敛证明,改进现有压缩性能
深度学习模型参数量持续增长,亟需高效压缩技术以部署于资源受限设备。本文从信息几何角度分析模型压缩中的算子分解方法,强调核心挑战在于定义最优低计算子流形并实现投影。我们指出,许多成功的压缩方法本质上是隐式近似该投影的信息散度。对于预训练模型,使用信息散度对提升零样本准确率至关重要;但在微调场景下,瓶颈模型的可训练性变得更为关键,因此需要采用迭代优化方法。本文证明了在软秩约束下,迭代奇异值阈值法在训练神经网络时具有收敛性。为进一步验证该视角的实用性,我们展示通过简化方法引入更柔和的秩降低策略,可在固定压缩率下实现性能提升。
原文摘要 · Abstract (English)
The ever-increasing parameter counts of deep learning models necessitate effective compression techniques for deployment on resource-constrained devices. This paper explores the application of information geometry, the study of density-induced metrics on parameter spaces, to analyze existing methods within the space of model compression, primarily focusing on operator factorization. Adopting this perspective highlights the core challenge: defining an optimal low-compute submanifold (or subset) and projecting onto it. We argue that many successful model compression approaches can be understood as implicitly approximating information divergences for this projection. We highlight that when compressing a pre-trained model, using information divergences is paramount for achieving improved zero-shot accuracy, yet this may no longer be the case when the model is fine-tuned. In such scenarios, trainability of bottlenecked models turns out to be far more important for achieving high compression ratios with minimal performance degradation, necessitating adoption of iterative methods. In this context, we prove convergence of iterative singular value thresholding for training neural networks subject to a soft rank constraint. To further illustrate the utility of this perspective, we showcase how simple modifications to existing methods through softer rank reduction result in improved performance under fixed compression rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。