arXiv:2507.06367cs.LGmath.AG2025-07被引 2

揭示线性卷积网络梯度流的黎曼几何结构,无需特殊初始化。

The Riemannian Geometry Associated to Gradient Flows of Linear Convolutional Networks

  • 将参数空间梯度流转化为函数空间的黎曼梯度流,适用于任意初始化。
  • 对二维及以上卷积及步幅大于一的一维卷积均成立,覆盖主流架构。
  • 为理解深度学习优化提供几何视角,适合研究优化与表示理论者。

我们研究了学习深度线性卷积网络时梯度流的几何特性。对于线性全连接网络,近期研究表明,若权重初始化满足平衡条件,其在参数空间的梯度流可表示为函数空间(即权矩阵乘积空间)上的黎曼梯度流。本文证明:线性卷积网络在参数空间的梯度流可始终表示为函数空间上的黎曼梯度流,且不依赖于初始化条件。该结论对 D ≥ 2 的多维卷积成立;对 D = 1 时,只要所有卷积步幅大于一亦成立。对应的黎曼度量依赖于初始值。

原文摘要 · Abstract (English)

We study geometric properties of the gradient flow for learning deep linear convolutional networks. For linear fully connected networks, it has been shown recently that the corresponding gradient flow on parameter space can be written as a Riemannian gradient flow on function space (i.e., on the product of weight matrices) if the initialization satisfies a so-called balancedness condition. We establish that the gradient flow on parameter space for learning linear convolutional networks can be written as a Riemannian gradient flow on function space regardless of the initialization. This result holds for $D$-dimensional convolutions with $D \geq 2$, and for $D =1$ it holds if all so-called strides of the convolutions are greater than one. The corresponding Riemannian metric depends on the initialization.

深度学习优化理论黎曼几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。