针对低秩训练的优化难题,提出几何感知的新优化方法。
A geometric framework for momentum-based optimizers for low-rank training
- 基于动态低秩逼近构建新优化器,显式考虑参数空间几何结构。
- 实验显示在相同参数量下收敛更快,验证指标更优。
- 适合需要高效训练大模型的研究者和工程人员。
低秩预训练与微调近年成为降低大模型计算与存储开销的有力手段。然而,传统优化器如动量法或Adam在处理权重低秩参数化时面临潜在困难。本文揭示经典动量方法因优化景观的几何特性,可能无法收敛至局部最优。为此,提出源自动态低秩逼近的新训练策略,结合动力学低秩逼近与动量优化工具,设计出尊重参数空间内在几何结构的优化器。数值实验验证了该方法在给定参数预算下实现更快收敛与更强验证性能。
原文摘要 · Abstract (English)
Low-rank pre-training and fine-tuning have recently emerged as promising techniques for reducing the computational and storage costs of large neural networks. Training low-rank parameterizations typically relies on conventional optimizers such as heavy ball momentum methods or Adam. In this work, we identify and analyze potential difficulties that these training methods encounter when used to train low-rank parameterizations of weights. In particular, we show that classical momentum methods can struggle to converge to a local optimum due to the geometry of the underlying optimization landscape. To address this, we introduce novel training strategies derived from dynamical low-rank approximation, which explicitly account for the underlying geometric structure. Our approach leverages and combines tools from dynamical low-rank approximation and momentum-based optimization to design optimizers that respect the intrinsic geometry of the parameter space. We validate our methods through numerical experiments, demonstrating faster convergence, and stronger validation metrics at given parameter budgets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。