提出理论框架,解释为何数据驱动的低秩压缩比无数据方法更优。
Theoretical Guarantees for Low-Rank Compression of Deep Neural Networks
- 基于激活值的低秩结构,构建数据驱动压缩的理论分析框架
- 在渐弱假设下证明三个恢复定理,揭示压缩性能边界
- 为高效推理提供理论支持,适合模型压缩研究者参考
深度神经网络在众多应用中达到顶尖性能,但其高内存与计算需求在资源受限环境中带来挑战。模型压缩技术如低秩近似,可通过降低网络规模与复杂度,在几乎不损失精度的前提下缓解问题。本文提出一种数据驱动的后训练低秩压缩分析框架。在对激活值近似低秩结构的不同假设下,我们证明了三个恢复定理,将建模偏差视为噪声。结果推动了对数据驱动低秩压缩优于数据无关方法的理论理解,并为可保证性能的压缩算法设计提供了基础,有效降低推理成本。
原文摘要 · Abstract (English)
Deep neural networks have achieved state-of-the-art performance across numerous applications, but their high memory and computational demands present significant challenges, particularly in resource-constrained environments. Model compression techniques, such as low-rank approximation, offer a promising solution by reducing the size and complexity of these networks while only minimally sacrificing accuracy. In this paper, we develop an analytical framework for data-driven post-training low-rank compression. We prove three recovery theorems under progressively weaker assumptions about the approximate low-rank structure of activations, modeling deviations via noise. Our results represent a step toward explaining why data-driven low-rank compression methods outperform data-agnostic approaches and towards theoretically grounded compression algorithms that reduce inference costs while maintaining performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。