arXiv:2410.08451math.NAcs.LG2024-10被引 3

Kolmogorov-Arnold定理揭示了神经网络分层学习的数学原理。

The Proof of Kolmogorov-Arnold May Illuminate Neural Network Learning

  • 用外微分形式解析网络层间映射的雅可比矩阵稀疏性
  • 发现高阶外积稀疏性与深层网络概念演进相关
  • 适合研究深度学习理论基础的学者参考

Kolmogorov和Arnold在解答希尔伯特第13问题(针对连续函数)时,为现代神经网络理论奠定了基础。他们的证明将多变量函数表示分为两步:第一(非线性)层间映射将数据流形普遍嵌入单个隐藏层,其像呈现特定模式,使得后续动态可求解第二层间映射。我将此模式解释为几乎处处定义的层间映射雅可比矩阵的“次集中”现象。次集中意味着雅可比矩阵的高阶外积具有稀疏性。我们提出一个概念性论证,说明这种稀疏性可能为当代深度神经网络中逐级更高阶概念的涌现创造条件,并建议两类实验来验证该假设。

原文摘要 · Abstract (English)

Kolmogorov and Arnold, in answering Hilbert's 13th problem (in the context of continuous functions), laid the foundations for the modern theory of Neural Networks (NNs). Their proof divides the representation of a multivariate function into two steps: The first (non-linear) inter-layer map gives a universal embedding of the data manifold into a single hidden layer whose image is patterned in such a way that a subsequent dynamic can then be defined to solve for the second inter-layer map. I interpret this pattern as "minor concentration" of the almost everywhere defined Jacobians of the interlayer map. Minor concentration amounts to sparsity for higher exterior powers of the Jacobians. We present a conceptual argument for how such sparsity may set the stage for the emergence of successively higher order concepts in today's deep NNs and suggest two classes of experiments to test this hypothesis.

神经网络深度学习数学理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。