arXiv:2606.05899cs.LGcond-mat.dis-nn2026-06被引 2

解析LoRA微调在注意力模型中的高维规律,揭示预训练与微调的深层关系。

High-Dimensional Theory of LoRA Fine-Tuning in a Solvable Attention Model

  • 构建可解的单头注意力模型,分析预训练后低秩微调过程。
  • 在高维极限下,测试误差与表征对齐均可由有限阶参数精确预测。
  • 发现测试误差与表征质量不一致现象,指导主动微调策略设计。

我们发展了一种关于注意力模型中低秩适配(LoRA)的高维统计理论,刻画了预训练与微调之间的相互作用。引入一个可解框架:先在数据丰富的任务上预训练单头注意力层,再在有限数据上通过秩一LoRA更新进行微调。在高维极限下,两个阶段均能以一组有限阶参数实现精确渐近表征,从而对测试误差和表征对齐给出明确预测。分析表明,预训练的影响可归纳为一个有效噪声项,由此推导出最优预训练方案。此外,我们揭示了一个测试误差与表征质量不匹配的区间,并提出将该理论应用于主动微调。

原文摘要 · Abstract (English)

We develop a high-dimensional statistical theory of low-rank adaptation (LoRA) in attention models, capturing the interplay between pre-training and fine-tuning. We introduce a solvable framework in which a single-head attention layer is first pre-trained on a data-abundant task and subsequently adapted via a rank-one LoRA update on limited data. In the high-dimensional limit, both stages admit a sharp asymptotic characterization in terms of a finite set of order parameters, yielding explicit predictions for test errors and representation alignment. Our analysis shows that the impact of pre-training on LoRA is summarized by an effective noise term, from which we derive prescriptions for the optimal pre-training procedure. We also demonstrate a regime with a mismatch between the value of the test error and representation quality, and propose an application of our theory to active fine-tuning.

LoRA注意力模型微调理论高维统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。