证明自监督网络在无限宽时呈现核行为,为理论分析提供新依据。
Infinite Width Limits of Self Supervised Neural Networks
- 基于巴洛双子损失,推导二层网络在无限宽下的核特性。
- 证明网络宽度趋于无穷时,其核函数保持恒定不变。
- 首次为自监督学习的核方法提供理论支撑,适合研究者参考。
NTK 是深度学习理论分析中广泛使用的工具,可将监督学习的深层神经网络通过核回归视角进行考察。近期多项研究探讨了自监督学习的核模型,假设这些模型也能揭示宽网络的行为,得益于 NTK。然而,这种联系在数学上是否成立仍存疑问——普遍存在的误解是,宽网络的核行为与训练损失函数无关。本文填补了这一空白,聚焦于在巴洛双子(Barlow Twins)损失下训练的两层神经网络,证明当网络宽度趋于无穷时,其 NTK 确实保持恒定。所采用的分析方法不同于以往工作,可能具有独立意义。整体而言,本研究首次为经典核理论应用于宽网络自监督学习提供了理论支持。基于此结果,我们推导出核化巴洛双子模型的泛化误差界,并将其与有限宽度神经网络相连接。
原文摘要 · Abstract (English)
The NTK is a widely used tool in the theoretical analysis of deep learning, allowing us to look at supervised deep neural networks through the lenses of kernel regression. Recently, several works have investigated kernel models for self-supervised learning, hypothesizing that these also shed light on the behavior of wide neural networks by virtue of the NTK. However, it remains an open question to what extent this connection is mathematically sound -- it is a commonly encountered misbelief that the kernel behavior of wide neural networks emerges irrespective of the loss function it is trained on. In this paper, we bridge the gap between the NTK and self-supervised learning, focusing on two-layer neural networks trained under the Barlow Twins loss. We prove that the NTK of Barlow Twins indeed becomes constant as the width of the network approaches infinity. Our analysis technique is a bit different from previous works on the NTK and may be of independent interest. Overall, our work provides a first justification for the use of classic kernel theory to understand self-supervised learning of wide neural networks. Building on this result, we derive generalization error bounds for kernelized Barlow Twins and connect them to neural networks of finite width.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。