揭示神经网络在连续流形上学习离散计算的几何机制
Emergent Riemannian geometry over learning discrete computations on continuous manifolds
- 通过黎曼拉回度量分析网络表征,发现计算可分解为特征离散化与逻辑运算
- 不同学习范式(丰富/懒惰)导致显著不同的度量与曲率结构
- 为理解神经网络如何处理连续输入的离散任务提供新几何视角
许多任务需要将连续输入数据(如图像)映射到离散输出(如类别标签)。然而,神经网络如何在连续数据流形上学习执行此类离散计算仍不清楚。本文通过分析神经网络各层间的黎曼拉回度量,发现网络计算可分解为两个函数:对连续输入特征进行离散化,以及在这些离散变量上执行逻辑操作。此外,我们展示了不同学习范式(丰富 vs. 懒惰)具有截然不同的度量和曲率结构,影响网络对未见输入的泛化能力。总体而言,本工作提出了一个几何框架,用以理解神经网络如何在连续流形上学习执行离散计算。
原文摘要 · Abstract (English)
Many tasks require mapping continuous input data (e.g. images) to discrete task outputs (e.g. class labels). Yet, how neural networks learn to perform such discrete computations on continuous data manifolds remains poorly understood. Here, we show that signatures of such computations emerge in the representational geometry of neural networks as they learn. By analysing the Riemannian pullback metric across layers of a neural network, we find that network computation can be decomposed into two functions: discretising continuous input features and performing logical operations on these discretised variables. Furthermore, we demonstrate how different learning regimes (rich vs. lazy) have contrasting metric and curvature structures, affecting the ability of the networks to generalise to unseen inputs. Overall, our work provides a geometric framework for understanding how neural networks learn to perform discrete computations on continuous manifolds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。