arXiv:2412.03035cs.LG2024-12被引 1

从因果视角重看梯度下降,揭示剪枝中的相变现象。

A Granger-Causal Perspective on Gradient Descent with Application to Pruning

  • 将梯度下降视为隐含因果过程,显式构建损失与参数变化的因果关系。
  • 剪枝比例增大时出现相变,暗示存在最优剪枝策略。
  • 剪枝后极小值更平坦,解释为何准确率反而提升,适合模型压缩研究者。

随机梯度下降(SGD)是优化神经网络的主要方法。深度网络的若干泛化特性,如收敛到更平坦的极小值,被认为源于SGD。本文从因果性角度探讨梯度下降过程,证明其隐含损失下降与参数更新之间的格兰杰因果关系。通过适当修改,使这一因果关系显式化。基于此因果框架,可实现更精细的控制。本文以剪枝为例展示其应用价值:(i)随着剪枝参数比例增加,观察到明显的相变现象,提示存在最优剪枝策略;(ii)剪枝后极小值变得更平坦,解释了剪枝后准确率提升的现象。

原文摘要 · Abstract (English)

Stochastic Gradient Descent (SGD) is the main approach to optimizing neural networks. Several generalization properties of deep networks, such as convergence to a flatter minima, are believed to arise from SGD. This article explores the causality aspect of gradient descent. Specifically, we show that the gradient descent procedure has an implicit granger-causal relationship between the reduction in loss and a change in parameters. By suitable modifications, we make this causal relationship explicit. A causal approach to gradient descent has many significant applications which allow greater control. In this article, we illustrate the significance of the causal approach using the application of Pruning. The causal approach to pruning has several interesting properties - (i) We observe a phase shift as the percentage of pruned parameters increase. Such phase shift is indicative of an optimal pruning strategy. (ii) After pruning, we see that minima becomes "flatter", explaining the increase in accuracy after pruning weights.

梯度下降剪枝因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。