揭示宽神经网络特征学习中的隐含正则化机制,统一了核与特征学习两种范式。
Canonical Regularisation of Wide Feature-Learning Neural Networks

- 提出泛化正则化能量函数,统一描述两类神经网络的训练偏好
- 发现即使在极小正则下,特征学习仍受岭回归偏置影响
- 引入弧岭作为可扩展鲁棒代理,关联早停与正则性本质
宽神经网络在特征学习范式中推动现代深度学习发展,但其研究远不如核范式充分。本文聚焦二者间关键却未被充分探讨的差异:梯度流训练所隐含的正则化与先验。该经典正则化性质在核范式中已有深入研究——梯度流在所有无穷全局最小解中精确选择零岭解,并支撑著名的神经网络高斯过程(NN-GP)对应关系,实现训练噪声建模。然而我们证明,在特征学习范式中,即使在正则化趋近于零的极限下,岭正则仍对梯度流产生偏置。训练过程中,岭正则扭曲了网络的归纳偏置,尤其对预训练网络造成显著影响,因其隐含先验具有信息量。为此,我们通过公理化将规范正则化定义为与范式无关的函数空间能量,并由此提升出一个唯一识别核范式中岭项的框架,且关键地推广至特征学习范式。通过研究特征学习网络的黎曼几何结构,我们推导出测地岭,实现了岭正则在特征学习范式的泛化。相应地,我们证明规范函数空间先验为黎曼吉布斯过程,是对更熟悉的高斯过程的推广。作为实用贡献,我们提出弧岭作为极小极大鲁棒、可扩展的测地岭替代方案,揭示早停与规范正则化在不同学习范式间的深层联系。最后,我们在图像处理和NLP迁移学习任务上实证验证了该理论的影响。
原文摘要 · Abstract (English)
Wide neural networks in the feature-learning regime drive modern deep learning, and yet they remain far less studied than their kernel-regime counterparts. We consider a critical yet under-explored difference between these two regimes: the regulariser and prior implied by gradient flow training. This canonical regularisation property is well-studied in kernel regime networks -- of all the infinite global minima, gradient flow selects exactly the vanishing ridge solution -- and underpins the celebrated NN-GP correspondence, precisely allowing the modelling of noise during training. However, we prove ridge regularisation biases gradient flow in feature-learning regime networks, even in the infinitesimal limit of vanishing regularisation. Over training, ridge distorts the inductive bias of the network, with a particular damage done to pretrained networks where the implicit prior is informative. We resolve this by axiomatising the canonical regulariser as a regime-agnostic function-space energy and lift, which uniquely identifies ridge in the kernel regime, and crucially generalises to the feature-learning regime. By studying the Riemannian geometry of feature-learning networks, we derive geodesic ridge from our framework, generalising ridge to the feature-learning regime. Correspondingly, we prove the canonical function-space prior is a Riemannian Gibbs Process, generalising the more familiar Gaussian Process. As a practical contribution, we propose arc ridge as a minimax-robust, scalable surrogate to geodesic ridge, revealing a deep relationship between early stopping and canonical regularisation across learning regimes. Finally, we demonstrate the consequences of our theory empirically on both image processing and NLP transfer-learning problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。