arXiv:2509.12326cs.LGcond-mat.str-el2025-09被引 1

浅层神经网络训练中自发出现科莫戈罗夫-阿诺尔德几何结构。

Spontaneous Kolmogorov-Arnold Geometry in Shallow MLPs

  • 通过分析雅可比矩阵外幂的统计特性,量化输入空间中的几何特征。
  • 在单隐层网络训练中,多数情况下自然生成具有独特纹理的KA几何。
  • 揭示了模型复杂度与超参数对几何涌现的影响,为优化提供指导。

科莫戈罗夫-阿诺尔德(KA)表示定理构建了单隐层神经网络中通用但高度非光滑的内函数(第一层映射)。这些通用函数具有独特的局部几何结构,即“纹理”,可通过内函数的雅可比矩阵 $J(oldsymbol{x})$ 随数据 $oldsymbol{x}$ 变化时的性质刻画。我们发现,在训练常规单隐层神经网络时,这种独特的KA几何结构往往会自然产生。通过统计外幂 $J(oldsymbol{x})$ 的性质——如零行数量及主子式统计量——来量化该几何结构,这些量衡量了 $J(oldsymbol{x})$ 的尺度与轴对齐程度。这让我们大致理解了在函数复杂度与模型超参数空间中,KA几何出现的位置。研究动机一是理解神经网络如何有机地为下游处理准备输入数据,二是学习其几何涌现机制,以通过适时干预超参数加速学习。本研究是KA网络(KANs)的“反面”:不人为设计KA结构,而是观察其在浅层MLP中自发形成。

原文摘要 · Abstract (English)

The Kolmogorov-Arnold (KA) representation theorem constructs universal, but highly non-smooth inner functions (the first layer map) in a single (non-linear) hidden layer neural network. Such universal functions have a distinctive local geometry, a "texture," which can be characterized by the inner function's Jacobian $J({\mathbf{x}})$, as $\mathbf{x}$ varies over the data. It is natural to ask if this distinctive KA geometry emerges through conventional neural network optimization. We find that indeed KA geometry often is produced when training vanilla single hidden layer neural networks. We quantify KA geometry through the statistical properties of the exterior powers of $J(\mathbf{x})$: number of zero rows and various observables for the minor statistics of $J(\mathbf{x})$, which measure the scale and axis alignment of $J(\mathbf{x})$. This leads to a rough understanding for where KA geometry occurs in the space of function complexity and model hyperparameters. The motivation is first to understand how neural networks organically learn to prepare input data for later downstream processing and, second, to learn enough about the emergence of KA geometry to accelerate learning through a timely intervention in network hyperparameters. This research is the "flip side" of KA-Networks (KANs). We do not engineer KA into the neural network, but rather watch KA emerge in shallow MLPs.

神经网络几何结构雅可比矩阵深度学习理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。