arXiv:2501.06074cs.LGmath.AG2025-01中稿 · SIAM Journal on Ap…被引 10

研究浅层多项式网络的几何与优化,揭示宽度与训练数据分布对收敛性的影响。

Geometry and Optimization of Shallow Polynomial Networks

  • 用对称张量建模单输出浅层网络,分析宽度与优化关系
  • 提出教师-度量判别器,刻画数据分布对优化行为的影响
  • 给出二次激活网络在高斯数据下的临界点完整分类

我们研究具有单项式激活函数且输出维度为一的浅层神经网络。这些模型的函数空间可识别为具有有界秩的对称张量集合。本文描述了这类网络的一般特性,重点关注宽度与优化之间的关系。随后考虑教师-学生问题,其可视为在非标准内积下进行低秩张量逼近的问题,该内积由数据分布诱导。在此框架中,我们引入教师-度量数据判别器,编码优化行为随训练数据分布变化的定性特征。最后聚焦于二次激活网络,深入分析其优化景观。特别地,我们提出了埃克哈特-杨定理的一个变体,刻画了在高斯训练数据下,教师-学生问题中所有临界点及其Hessian符号特征。

原文摘要 · Abstract (English)

We study shallow neural networks with monomial activations and output dimension one. The function space for these models can be identified with a set of symmetric tensors with bounded rank. We describe general features of these networks, focusing on the relationship between width and optimization. We then consider teacher-student problems, which can be viewed as problems of low-rank tensor approximation with respect to non-standard inner products that are induced by the data distribution. In this setting, we introduce a teacher-metric data discriminant which encodes the qualitative behavior of the optimization as a function of the training data distribution. Finally, we focus on networks with quadratic activations, presenting an in-depth analysis of the optimization landscape. In particular, we present a variation of the Eckart-Young Theorem characterizing all critical points and their Hessian signatures for teacher-student problems with quadratic networks and Gaussian training data.

神经网络张量分析优化理论深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。