arXiv:2503.08502cs.LGcs.NE2025-03中稿 · ICLR

通过量化神经网络的折叠程度,提升模型泛化能力。

The Space Between: On Folding, Symmetries and Sampling

  • 提出基于等价类的新折叠度量方法,适用于多种激活函数。
  • 深度越深折叠值越高,且与低泛化误差正相关。
  • 设计新正则化策略,引导网络学习更高折叠解。

近期研究表明,使用ReLU激活函数的神经网络在训练过程中会折叠输入空间。尽管已有诸多工作暗示此现象,但直到最近才出现基于ReLU激活空间汉明距离的折叠度量方法。本文将该度量推广至更广泛的激活函数类别,引入输入数据的等价类概念,分析其数学与计算性质,并提出高效的采样实现策略。此外,观察发现当泛化误差较低时,网络深度增加会导致折叠值上升;而当误差增大时,折叠值下降。这表明数据流形中的对称性(如反射不变性)会在空间折叠中显现,从而增强网络泛化能力。受此启发,本文提出一种新型正则化方案,鼓励网络寻找具有更高折叠值的解。

原文摘要 · Abstract (English)

Recent findings suggest that consecutive layers of neural networks with the ReLU activation function \emph{fold} the input space during the learning process. While many works hint at this phenomenon, an approach to quantify the folding was only recently proposed by means of a space folding measure based on Hamming distance in the ReLU activation space. We generalize this measure to a wider class of activation functions through introduction of equivalence classes of input data, analyse its mathematical and computational properties and come up with an efficient sampling strategy for its implementation. Moreover, it has been observed that space folding values increase with network depth when the generalization error is low, but decrease when the error increases. This underpins that learned symmetries in the data manifold (e.g., invariance under reflection) become visible in terms of space folds, contributing to the network's generalization capacity. Inspired by these findings, we outline a novel regularization scheme that encourages the network to seek solutions characterized by higher folding values.

神经网络折叠现象正则化泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。