揭示深度单项网络中简化结构的数学根源,解释为何模型偏好简单函数。
Singular Learning and Occam's Razor in Deep Monomial Networks
- 用多项式代数分析深度全连接网络的临界点
- 发现高阶激活下临界点对应于部分神经元失效的子网络
- 为模型隐式偏向简单函数提供理论解释,适合学习理论者
在神经网络优化中,梯度动态受模型架构引发的临界点影响。这些临界点出现在参数化雅可比矩阵秩亏时,是奇异学习理论研究的核心奇点。本文通过多项式代数工具(如Mason定理)研究具有单项激活函数的深层全连接网络中的此类点。结果表明,在足够大的激活阶数下,临界性恰好出现在子网络上,即某些神经元处于非活跃或冗余状态的参数配置。这一发现从数学上揭示了深度神经网络的隐式偏差,解释了模型倾向于收敛到更简单函数的现象。
原文摘要 · Abstract (English)
In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture. These critical points occur where the Jacobian of the model's parametrization is rank-deficient, and are the most pronounced singularities studied in Singular Learning Theory. We investigate such points in deep fully-connected networks with monomial activations via tools from polynomial algebra such as Mason's Theorem. We show that, for sufficiently large activation degree, criticality occurs precisely at subnetworks, i.e., at parameter configurations where some neurons are inactive or redundant. This offers a mathematical perspective on the implicit bias in deep neural networks, explaining the tendency of these models to converge toward simpler functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。