arXiv:2603.03610cs.LG2026-03

用黎曼几何重新理解反向传播,提升模块化系统的优化效率。

Riemannian Optimization in Modular Systems

  • 将反向传播视为约束优化问题,结合黎曼梯度轨迹的最小作用量原理。
  • 提出分层黎曼度量,利用Woodbury恒等式实现高效计算,避免$O(n^3)$开销。
  • 构建可组合的黎曼模块框架,提供$O(κ^2 L/(ξμ oot{n})$的收敛稳定性保障。

理解由模块化组件构成的系统如何协同优化,是生物学、工程学和机器学习中的关键问题。反向传播算法是其中一种解决方案,对神经网络的成功至关重要。尽管其在实践中表现优异,但理论基础仍不充分。本文结合黎曼几何、最优控制理论与理论物理工具,推进对此的理解。主要贡献有三:首先,将反向传播重新推导为约束优化问题,并结合黎曼梯度下降轨迹是最小作用量路径的洞察;其次,引入递归定义的分层黎曼度量,利用Woodbury矩阵恒等式高效计算,避免全度量求逆的$O(n^3)$代价;第三,构建可组合的“黎曼模块”框架,基于非线性收缩理论量化收敛性,提供$O(κ^2 L/(ξμ oot{n}))$阶的算法稳定性保证,其中$κ$和$L$为Lipschitz常数,$μ$为质量矩阵尺度,$ξ$界定了条件数。该分层度量方法为自然梯度下降提供了实用替代方案。虽聚焦于神经网络,本方法更普遍适用于随时间优化的模块化系统,如生物进化与发育过程。

原文摘要 · Abstract (English)

Understanding how systems built out of modular components can be jointly optimized is an important problem in biology, engineering, and machine learning. The backpropagation algorithm is one such solution and has been instrumental in the success of neural networks. Despite its empirical success, a strong theoretical understanding of it is lacking. Here, we combine tools from Riemannian geometry, optimal control theory, and theoretical physics to advance this understanding. We make three key contributions: First, we revisit the derivation of backpropagation as a constrained optimization problem and combine it with the insight that Riemannian gradient descent trajectories can be understood as the minimum of an action. Second, we introduce a recursively defined layerwise Riemannian metric that exploits the modular structure of neural networks and can be efficiently computed using the Woodbury matrix identity, avoiding the $O(n^3)$ cost of full metric inversion. Third, we develop a framework of composable ``Riemannian modules'' whose convergence properties can be quantified using nonlinear contraction theory, providing algorithmic stability guarantees of order $O(κ^2 L/(ξμ\sqrt{n}))$ where $κ$ and $L$ are Lipschitz constants, $μ$ is the mass matrix scale, and $ξ$ bounds the condition number. Our layerwise metric approach provides a practical alternative to natural gradient descent. While we focus here on studying neural networks, our approach more generally applies to the study of systems made of modules that are optimized over time, as it occurs in biology during both evolution and development.

黎曼优化反向传播模块化系统稳定保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。