arXiv:2412.06545cs.LG2024-12被引 2

发现迭代剪枝通过增强非高斯性,让全连接网络自动形成局部感受野。

On How Iterative Magnitude Pruning Discovers Local Receptive Fields in Fully Connected Neural Networks

  • 用新方法测量权重对表示统计特性的影响
  • 剪枝过程逐步提升前激活的非高斯性
  • 适合研究模型归纳偏置与压缩机制的人

自彩票理论提出以来,迭代幅度剪枝(IMP)已成为提取高性能稀疏子网络的常用方法。尽管效果显著,其内在机制仍不明确。有研究指出,将IMP应用于全连接神经网络(FCNs)会引发局部感受野(RFs)的出现,这与哺乳动物视觉皮层和卷积网络的特性一致。然而,为何剪枝能发现局部特征尚不清楚。受启发于合成图像中非高斯统计(如锐边)可驱动FCN产生局部RFs的现象,我们假设IMP通过迭代增加FCN表示的非高斯性,形成反馈循环以增强局部化。本文首次证明,非高斯输入是诱发局部RFs的必要条件;并提出一种新方法(腔方法)量化单个权重对表示统计的影响,证实了IMP系统性地提升前激活的非高斯性,从而促成局部感受野的形成。本工作为理解IMP如何生成强归纳偏置提供了简洁解释。

原文摘要 · Abstract (English)

Since its use in the Lottery Ticket Hypothesis, iterative magnitude pruning (IMP) has become a popular method for extracting sparse subnetworks that can be trained to high performance. Despite its success, the mechanism that drives the success of IMP remains unclear. One possibility is that IMP is capable of extracting subnetworks with good inductive biases that facilitate performance. Supporting this idea, recent work showed that applying IMP to fully connected neural networks (FCNs) leads to the emergence of local receptive fields (RFs), a feature of mammalian visual cortex and convolutional neural networks that facilitates image processing. However, it remains unclear why IMP would uncover localized features in the first place. Inspired by results showing that training on synthetic images with highly non-Gaussian statistics (e.g., sharp edges) is sufficient to drive the emergence of local RFs in FCNs, we hypothesize that IMP iteratively increases the non-Gaussian statistics of FCN representations, creating a feedback loop that enhances localization. Here, we demonstrate first that non-Gaussian input statistics are indeed necessary for IMP to discover localized RFs. We then develop a new method for measuring the effect of individual weights on the statistics of the FCN representations ("cavity method"), which allows us to show that IMP systematically increases the non-Gaussianity of pre-activations, leading to the formation of localized RFs. Our work, which is the first to study the effect of IMP on the statistics of the representations of neural networks, sheds parsimonious light on one way in which IMP can drive the formation of strong inductive biases.

剪枝归纳偏置非高斯性感受野

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。