提出新方法减少低秩训练中的QR分解次数,提升稳定性与效率
An Augmented Backward-Corrected Projector Splitting Integrator for Dynamical Low-Rank Training
- 在投影分裂法中引入增强步骤,降低对QR分解的依赖
- 理论证明可收敛至局部最优解,多基准测试验证有效性
- 适合需要高效低秩训练的模型优化场景
层分解已成为训练内存高效神经网络的常用技术。然而,层分解方法在训练过程中存在鲁棒性不足的问题。为克服这一局限,动态低秩训练方法被提出,利用稳健的时间积分技术求解低秩矩阵微分方程。尽管这些方法实现了高效训练,仍依赖于小秩矩阵的计算密集型QR分解和奇异值分解。本文提出一种新型低秩训练方法,显著减少所需QR分解次数。该方法将增强步骤融入投影分裂框架,确保收敛至局部最优解。我们提供了严谨的理论分析,并在多个基准上验证了其有效性。
原文摘要 · Abstract (English)
Layer factorization has emerged as a widely used technique for training memory-efficient neural networks. However, layer factorization methods face several challenges, particularly a lack of robustness during the training process. To overcome this limitation, dynamical low-rank training methods have been developed, utilizing robust time integration techniques for low-rank matrix differential equations. Although these approaches facilitate efficient training, they still depend on computationally intensive QR and singular value decompositions of matrices with small rank. In this work, we introduce a novel low-rank training method that reduces the number of required QR decompositions. Our approach integrates an augmentation step into a projector-splitting scheme, ensuring convergence to a locally optimal solution. We provide a rigorous theoretical analysis of the proposed method and demonstrate its effectiveness across multiple benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。