arXiv:2505.23869stat.MLcs.LG2025-05

压缩越高效,模型随机性越强,二者存在可计算的内在联系。

Gibbs randomness-compression proposition

  • 用吉布斯熵衡量压缩后权重向量的随机性,以任务性能为压缩探针。
  • 压缩比微降时,性能与熵呈共单调关系,验证了该关联的普适性。
  • 适用于研究模型压缩与泛化能力关系的研究者,尤其关注深层网络优化。

本文通过在损失性压缩过程相关的测量向量集上定义吉布斯熵,提出一个连接随机性与压缩性的命题。通过学习任务性能作为迭代压缩-训练循环中的压缩探针,从统计力学视角可视为热力学效率驱动的迭代粗粒化过程。我们以极小的压缩比下降为前提,建立压缩性能与熵之间的共单调关系。通过深度学习中经典的视觉任务,使用三种不同复杂度的模型压缩方法(基准模型)验证该命题:(1) 随机剪枝,(2) 基于幅度的剪枝,(3) 双重全息压缩(一种新型压缩感知双路应用方法)。将深度网络剩余权重作为测量向量,计算其吉布斯熵。结果表明,压缩性能与模型熵之间存在固有的、可计算的内在联系。

原文摘要 · Abstract (English)

A proposition that connects randomness and compression is put forward via Gibbs entropy over set of measurement vectors associated with a lossy compression process. In building this connection, we use a performance of a learning task as a probe of compression in iterative compress-train cycles. This can be thought as iterative coarse-graining from statistical mechanics perspective using thermodynamic efficiency as a probe. We formulate this connection via comonotonic relationship within a very small decrease in compression ratio and the performance. We have showcase the validity of this proposition with a canonical vision task in deep learning with three different model compression processes as {\it a baseline model}. We use the following, simpler to more complex model compression approaches: (1) random pruning,(2) magnitude pruning, and (3) a more complex compression by using dual tomographic compression, which utilizes compressed sensing in dual fashion which is introduced as a new method. We use remaining weights of deep learning network as a measurement vector where we measure the Gibbs entropy. We show case the idea that there is an inherent computable connection between compression probed by performance and randomness from an entropy measure on the learned model.

压缩深度学习随机性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。