arXiv:2502.02766cs.LGcs.IT2025-02被引 5

提出理论框架,解释为何数据驱动的低秩压缩比无数据方法更优。

Theoretical Guarantees for Low-Rank Compression of Deep Neural Networks

  • 基于激活值的低秩结构,构建数据驱动压缩的理论分析框架
  • 在渐弱假设下证明三个恢复定理,揭示压缩性能边界
  • 为高效推理提供理论支持,适合模型压缩研究者参考

深度神经网络在众多应用中达到顶尖性能,但其高内存与计算需求在资源受限环境中带来挑战。模型压缩技术如低秩近似,可通过降低网络规模与复杂度,在几乎不损失精度的前提下缓解问题。本文提出一种数据驱动的后训练低秩压缩分析框架。在对激活值近似低秩结构的不同假设下,我们证明了三个恢复定理,将建模偏差视为噪声。结果推动了对数据驱动低秩压缩优于数据无关方法的理论理解,并为可保证性能的压缩算法设计提供了基础,有效降低推理成本。

原文摘要 · Abstract (English)

Deep neural networks have achieved state-of-the-art performance across numerous applications, but their high memory and computational demands present significant challenges, particularly in resource-constrained environments. Model compression techniques, such as low-rank approximation, offer a promising solution by reducing the size and complexity of these networks while only minimally sacrificing accuracy. In this paper, we develop an analytical framework for data-driven post-training low-rank compression. We prove three recovery theorems under progressively weaker assumptions about the approximate low-rank structure of activations, modeling deviations via noise. Our results represent a step toward explaining why data-driven low-rank compression methods outperform data-agnostic approaches and towards theoretically grounded compression algorithms that reduce inference costs while maintaining performance.

模型压缩低秩近似理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。