预训练过强反而拖慢LoRA微调,理论揭示其动态机制
When pre-training hurts LoRA fine-tuning: a dynamical analysis via single-index models
- 用单指标模型分析LoRA微调动力学,解析预训练强度影响
- 强预训练导致搜索期延长,收敛速度下降,即使任务对齐也如此
- 理论适用于真实视觉模型,为微调策略提供新视角
预训练通常被认为有助于下游任务微调。本文通过数学分析发现,这一直觉并不总成立:过度预训练可能在计算上减缓微调优化过程。研究聚焦于单指标模型下的一次性随机梯度下降(one-pass SGD)训练中的低秩适配(LoRA)微调,利用微调动态的统计量描述,精确刻画了收敛速率如何依赖于初始对齐程度和目标任务的非线性程度。关键结论是,即便预训练与下游任务高度对齐,强烈的预训练仍会引发漫长的搜索阶段,阻碍收敛。该理论为预训练强度与任务难度共同作用下LoRA微调的动力学及局限性提供了统一图景。实证表明,这些理论发现可推广至真实数据上的视觉变换器模型。
原文摘要 · Abstract (English)
Pre-training on a source task is usually expected to facilitate fine-tuning on similar downstream problems. In this work, we mathematically show that this naive intuition is not always true: excessive pre-training can computationally slow down fine-tuning optimization. We study this phenomenon for low-rank adaptation (LoRA) fine-tuning on single-index models trained under one-pass SGD. Leveraging a summary statistics description of the fine-tuning dynamics, we precisely characterize how the convergence rate depends on the initial fine-tuning alignment and the degree of non-linearity of the target task. The key take away is that even when the pre-training and downstream tasks are well aligned, strong pre-training can induce a prolonged search phase and hinder convergence. Our theory thus provides a unified picture of how pre-training strength and task difficulty jointly shape the dynamics and limitations of LoRA fine-tuning in a nontrivial tractable model. On the practical side, we empirically show that our theoretical findings extend beyond our toy model and remain relevant in the context of a vision-transformer model trained on real data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。