提出无需投影和调参的流形优化方法,加速大模型低秩微调。
Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning

- 直接落地流形,避免昂贵正交化与参数调优。
- 理论证明全局收敛,迭代复杂度达最优水平。
- 应用于LoRA微调,训练更快且下游性能强。
流形上的优化在机器学习中具有重要意义。现有方法或依赖需高成本正交化的投影算子,或依赖精细的步长与惩罚参数调整。为此,本文提出一种无需投影且免惩罚参数的算法,直接将更新结果落在流形上。通过利用二次惩罚函数的强凸类似性质及流形的近端光滑性,建立了在固定与递减步长下的全局收敛性,并达到最优已知的迭代复杂度。随后,将大语言模型的低秩适配(LoRA)微调问题重新表述为流形优化问题,提出几何加速的Manifold-LoRA方法。该方法采用所提落地技术与精心设计的步长策略,显著加速训练过程。在基准数据集上的数值实验表明,该方法效率高且下游性能优异。
原文摘要 · Abstract (English)
Optimization over the Stiefel manifold plays a significant role in various machine learning tasks. Existing methods either use the retraction operators, requiring costly orthonormalization for large-scale matrices, or employ landing methods that rely on careful step size selection and penalty parameter tuning. To address these challenges, we propose a retraction-free and penalty parameter-free algorithm that directly lands on the manifold. By leveraging the strongly-convex-like property of the quadratic penalty function and the proximal smoothness of the Stiefel manifold, we establish global convergence guarantees with the best-known iteration complexities under both constant and diminishing step sizes. Then, we reformulate the low-rank adaptation (LoRA) fine-tuning problem for large language models as a manifold optimization problem, introducing Manifold-LoRA for geometry-accelerated adaptation. This approach employs the proposed landing technique and a carefully designed step size strategy to accelerate the training process. Numerical experiments on benchmark datasets demonstrate the efficiency and strong downstream performance of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。