arXiv:2508.03820cs.LGmath.OC2025-08被引 2

提出随机低秩微调理论框架,统一多种高效微调方法

Bernoulli-LoRA: A Theoretical Framework for Randomized Low-Rank Adaptation

  • 用伯努利概率机制随机选择更新矩阵,统一现有低秩微调策略
  • 证明了7种变体在非凸优化下的收敛性,涵盖梯度下降到自适应算法
  • 理论与实验结合,为高效微调提供可解释的数学基础,适合算法研究者

参数高效微调(PEFT)已成为适配大型基础模型的关键方法,尤其在模型规模持续指数增长的背景下。其中,低秩微调(LoRA)因高效简洁而突出,将调整表示为两个低秩矩阵的乘积。尽管大量实证研究验证了其有效性,但理论理解仍不充分。近期工作对RAC-LoRA进行了初步严谨分析。本文提出伯努利-LoRA,一种新颖的理论框架,统一并扩展了现有LoRA方法。该方法引入基于伯努利分布的概率机制,决定更新哪个矩阵,涵盖且泛化多种已有更新策略,同时保持理论可处理性。在非凸优化标准假设下,我们分析了多种变体:伯努利-LoRA-GD、伯努利-LoRA-SGD、伯努利-LoRA-PAGE、伯努利-LoRA-MVR、伯努利-LoRA-QGD、伯努利-LoRA-MARINA 和伯努利-LoRA-EF21,为每种变体建立了收敛性保证。此外,我们将分析扩展至凸非光滑函数,给出了常数步长和自适应(Polyak型)步长的收敛速率。通过在多种任务上的大量实验,我们验证了理论发现,并展示了方法的实际有效性。本工作是构建理论坚实且实用高效的PEFT方法的重要一步。

原文摘要 · Abstract (English)

Parameter-efficient fine-tuning (PEFT) has emerged as a crucial approach for adapting large foundational models to specific tasks, particularly as model sizes continue to grow exponentially. Among PEFT methods, Low-Rank Adaptation (LoRA) (arXiv:2106.09685) stands out for its effectiveness and simplicity, expressing adaptations as a product of two low-rank matrices. While extensive empirical studies demonstrate LoRA's practical utility, theoretical understanding of such methods remains limited. Recent work on RAC-LoRA (arXiv:2410.08305) took initial steps toward rigorous analysis. In this work, we introduce Bernoulli-LoRA, a novel theoretical framework that unifies and extends existing LoRA approaches. Our method introduces a probabilistic Bernoulli mechanism for selecting which matrix to update. This approach encompasses and generalizes various existing update strategies while maintaining theoretical tractability. Under standard assumptions from non-convex optimization literature, we analyze several variants of our framework: Bernoulli-LoRA-GD, Bernoulli-LoRA-SGD, Bernoulli-LoRA-PAGE, Bernoulli-LoRA-MVR, Bernoulli-LoRA-QGD, Bernoulli-LoRA-MARINA, and Bernoulli-LoRA-EF21, establishing convergence guarantees for each variant. Additionally, we extend our analysis to convex non-smooth functions, providing convergence rates for both constant and adaptive (Polyak-type) stepsizes. Through extensive experiments on various tasks, we validate our theoretical findings and demonstrate the practical efficacy of our approach. This work is a step toward developing theoretically grounded yet practically effective PEFT methods.

低秩微调理论分析参数高效优化理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。