解析LoRA的秩与精度关系,揭示低秩更新如何影响模型性能。
Rank-Accuracy Trade-off for LoRA: A Gradient-Flow Analysis
- 从动力系统角度分析LoRA梯度流,建立秩与精度的数学关系。
- 证明秩为1时仍可接近全参数精度,且在两种损失下有闭式解。
- 适用于研究模型压缩、高效微调的算法设计者与研究者。
先前的实证研究显示,即使使用秩为1的更新,LoRA在下游微调任务中也能达到与全参数方法相当的精度。然而,LoRA精度对更新秩的依赖关系的理论基础仍不清晰。本文从动力系统视角,对比了低秩(秩-r)与全参数更新在微调任务中的精度表现。我们在全秩与低秩场景下进行梯度流分析,推导出在两种损失函数(迹平方和Frobenius范数低秩逼近损失)下,LoRA秩与精度之间的显式数学关系。尽管已有工作提出过LoRA的梯度流方程,本文首次严格推导其形式,并证明同时与顺序更新下的方程一致。基于所得的动力系统方程,我们获得了两种损失函数下的闭式精度-秩关系。
原文摘要 · Abstract (English)
Previous empirical studies have shown that LoRA achieves accuracy comparable to full-parameter methods on downstream fine-tuning tasks, even for rank-1 updates. By contrast, the theoretical underpinnings of the dependence of LoRA's accuracy on update rank remain relatively unexplored. In this work, we compare the accuracy of rank-r LoRA updates against full-parameter updates for fine-tuning tasks from a dynamical systems perspective. We perform gradient flow analysis in both full-rank and low-rank regimes to establish explicit relationships between rank and accuracy for two loss functions under LoRA. While gradient flow equations for LoRA are presented in prior work, we rigorously derive their form and show that they are identical for simultaneous and sequential LoRA parameter updates. We then use the resulting dynamical system equations to obtain closed-form relationships between LoRA rank and accuracy for trace-squared and Frobenius-norm low-rank approximation loss functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。