揭示了低秩适配器在单次训练中的泛化能力边界,为高效微调提供理论支撑。
Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters
- 针对单次微调场景,分析随机低秩适配器的泛化性能集中性。
- 提出上界样本复杂度为 $ ilde{ m O}(rac{ ext{rank}^{1/2}}{ ext{样本数}^{1/2}})$,下界为 $ m O(rac{1}{ ext{样本数}^{1/2}})$。
- 适用于关注LoRA理论可靠性与参数效率的研究者或实践者。
低秩适配(LoRA)作为基础模型中广泛使用的参数高效微调(PEFT)技术,其低秩因子初始化存在固有的不对称性,该现象自提出以来一直未被理论解释。本文聚焦于冻结随机因子的非对称LoRA,首次在单次微调运行的框架下,刻画其泛化误差的集中特性。主结果表明,在 $N$ 个样本上训练秩为 $r$ 的LoRA时,以高概率成立的泛化误差上界为 $ ilde{ m O}(rac{ ext{rank}^{1/2}}{ ext{样本数}^{1/2}})$。同时,我们建立了匹配的下界 $ m O(rac{1}{ ext{样本数}^{1/2}})$,揭示了样本效率的根本极限。这些结果更贴近实际微调场景,为非对称LoRA的可靠性与实用性提供了关键理论依据。
原文摘要 · Abstract (English)
Low-Rank Adaptation (LoRA) has emerged as a widely adopted parameter-efficient fine-tuning (PEFT) technique for foundation models. Recent work has highlighted an inherent asymmetry in the initialization of LoRA's low-rank factors, which has been present since its inception and was presumably derived experimentally. This paper focuses on providing a comprehensive theoretical characterization of asymmetric LoRA with frozen random factors. First, while existing research provides upper-bound generalization guarantees based on averages over multiple experiments, the behaviour of a single fine-tuning run with specific random factors remains an open question. We address this by investigating the concentration of the typical LoRA generalization gap around its mean. Our main upper bound reveals a sample complexity of $\tilde{\mathcal{O}}\left(\frac{\sqrt{r}}{\sqrt{N}}\right)$ with high probability for rank $r$ LoRAs trained on $N$ samples. Additionally, we also determine the fundamental limits in terms of sample efficiency, establishing a matching lower bound of $\mathcal{O}\left(\frac{1}{\sqrt{N}}\right)$. By more closely reflecting the practical scenario of a single fine-tuning run, our findings offer crucial insights into the reliability and practicality of asymmetric LoRA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。