arXiv:2504.05279cs.LGhep-th2025-04被引 3

提出协变梯度下降,让优化在任意坐标系下保持一致。

Covariant Gradient Descent

  • 用梯度一阶二阶矩构造协变力与度量张量,实现坐标无关优化。
  • 计算复杂度保持线性,通过指数加权时间平均估计统计量。
  • 揭示RMSProp、Adam等为特例,可系统性改进现有优化器。

我们提出一种显式协变的梯度下降方法,确保在任意坐标系及一般弯曲可训练空间中的一致性。优化动态由协变力向量和协变度量张量定义,二者均基于梯度的一阶与二阶统计矩计算。这些矩通过带指数权重的时间平均进行估计,保持线性计算复杂度。我们证明,RMSProp、Adam 和 AdaBelief 等常用优化方法均为协变梯度下降(CGD)的特例,并展示了如何进一步推广与改进这些方法。

原文摘要 · Abstract (English)

We present a manifestly covariant formulation of the gradient descent method, ensuring consistency across arbitrary coordinate systems and general curved trainable spaces. The optimization dynamics is defined using a covariant force vector and a covariant metric tensor, both computed from the first and second statistical moments of the gradients. These moments are estimated through time-averaging with an exponential weight function, which preserves linear computational complexity. We show that commonly used optimization methods such as RMSProp, Adam and AdaBelief correspond to special limits of the covariant gradient descent (CGD) and demonstrate how these methods can be further generalized and improved.

优化算法协变学习深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。