用凯勒几何分析复杂神经网络优化,揭示负曲率对损失景观的破坏性影响。
Kähler landscapes for complex neural network descents and guarantees including a search and destroy of the Calabi-Yau manifold

- 引入凯勒信息度量,结合威廷格海森堡矩阵实现自然梯度下降。
- 在非紧凯勒流形上,常数行列式条件导致低秩度量引发爆炸效应。
- 揭示负截面曲率破坏优化保证,关联到神经网络训练失败机制。
我们研究复参数神经网络的优化景观。方法基于信息论流形视角与经典优化保证,结合多贝尔-奥斯特渐近分析等复几何工具。通过交叉熵定义下的威廷格海森堡矩阵,得到凯勒信息度量,使下降路径保持在全纯切丛内。采用自然梯度更新规则,其损失梯度经度量逆缩放。重点考察卡比拉-丘信息流形,该流形因曲率病态而缺乏理论保证。在非紧设定下,全局势能由几何定义而非依赖卡比拉猜想拓扑要求,楔积全纯形式为凯勒形式的外幂乘以常数,导出常数行列式条件。固定行列式时,若度量几乎低秩(满足特征值容差),则出现爆破现象。此外,发现负曲率会破坏损失景观,尤其截面曲率;进一步拓展至负定里奇曲率的关联。论证主要基于几何分析,但与深度学习理论建立联系,如初始化处渐近行为及零与负里奇曲率下模型保证失效的故障模式。
原文摘要 · Abstract (English)
We study landscapes for complex-parameterized networks. Our approach is motivated with an information-theoretic manifold perspective of the parameter and via classical optimization guarantees although of complex geometric variety such as through Dolbeault asymptotics. The descent path admits a Kähler information metric under a cross-entropy via the Wirtinger Hessian on the log-likelihood potential. We restrict attention to a descent update rule with natural gradient descent via a differentiated loss scaled by the inverse metric, so the descent path remains in the holomorphic tangent bundle. We emphasize Calabi-Yau information manifolds which profane theoretical guarantees via an ill-curvature-conditioned landscape. Under a Calabi-Yau metric, specifically in a non-compact setting with a global potential so defined geometrically rather than invoking the topological requirements of the Calabi conjecture, a wedged nowhere-vanishing holomorphic form is the top exterior product of the Kähler form up to constants, yielding a constant determinant condition. Under a fixed determinant, a metric almost low rank up to an eigenvalue tolerance implies a blow-up effect. Moreover, it has been discovered that negative curvature subverts the loss landscape, specifically sectional curvature, so we expand on this and draw interconnections to negative-definite Ricci curvature. Our arguments primarily exist in a geometric analytic modality, although we establish roots in deep learning theory such as through asymptotics at initialization and connections through failure modes of neural network guarantees under vanishing and negative Ricci curvature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。