提出新优化方法,让机器学习模型在数据曲面上更高效收敛。
Iso-Riemannian Optimization on Learned Data Manifolds
- 用保持欧氏速度的测地线结合投影梯度,改进传统优化方向。
- 理论证明该方法在特定凸性条件下可保证收敛,优于经典方法。
- 适合处理高维数据流形上的优化问题,尤其适用于无显式目标函数场景。
我们为约束于学习到的数据流形的优化问题构建了iso-Riemannian优化理论。传统黎曼优化(尤其是黎曼梯度下降)在此类问题中表现不佳:目标函数的优良欧氏性质未必带来测地凸性或L-光滑性,且黎曼梯度可能提供不合适的搜索方向。为此,我们结合由iso-connection诱导的流形映射(其测地线具有恒定欧氏速度)与l2投影的欧氏梯度,提出l2投影梯度iso-Riemannian下降法,以缓解上述问题。从两个互补视角分析该方案:首先引入iso-g-凸性和iso-L-光滑性,将其与Polyak-Lojasiewicz型条件关联,并建立通用收敛性保证,解决条件差与搜索方向不当的问题;其次引入向量场的iso-单调性和iso-Lipschitz性,虽在一维情形给出类似收敛结果并支持计算iso-Riemannian重心等应用,但在高维下无法自然推广。因此,在iso-Riemannian框架下,函数与向量场视角不再等价。这表明,不同经典黎曼优化的延伸导致不同假设与收敛保证。本文理论表明,对于函数优化,函数视角更具普适性、可操作性与适用性,而向量场视角仅适用于无直接函数形式的问题。
原文摘要 · Abstract (English)
We develop a theory of iso-Riemannian optimization for problems constrained to learned data manifolds, a setting in which classical Riemannian optimization - and Riemannian gradient descent in particular - can be poorly suited. That is, favorable Euclidean properties of an objective need not translate into geodesic convexity or L-smoothness, and the Riemannian gradient may provide an unsuitable search direction. We instead combine the manifold mappings induced by the iso-connection, whose geodesics have constant Euclidean speed, with the l2-projected Euclidean gradient, resulting in l2-projected gradient iso-Riemannian descent, in an attempt to alleviate these issues. We analyze this scheme from two complementary perspectives. First, we introduce iso-g-convexity and iso-L-smoothness, relate strong iso-g-convexity to a Polyak-Lojasiewicz-type condition, and establish general convergence guarantees. This function-based theory addresses the conditioning and search-direction issues that motivate our approach. Second, we introduce iso-monotonicity and iso-Lipschitzness for vector fields. However, while these notions yield analogous convergence results in one dimension and enable applications such as the computation of iso-Riemannian barycentres, the resulting theory need not extend meaningfully to higher dimensions. So unlike in the classical Levi-Civita setting, these function- and vector field-based perspectives need not be equivalent in the iso-Riemannian setting. Consequently, distinct extensions of classical Riemannian optimization lead to different assumptions and convergence guarantees. The theory developed here suggests that, for function optimization, the function-based perspective provides the more general, tractable and overall suitable framework, whereas the vector field perspective is naturally reserved for problems without a direct function-based formulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。