arXiv:2607.20301cs.LGcs.CL2026-07

高维空间近正交性解释了模型微调后长期有效的原因

The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability

  • 利用高维向量近正交性,解释为何微调参数可长期复用
  • 在10轮持续预训练中,微调效果仍保持稳定
  • 适合关注模型高效适配与理论机制的研究者

微调广泛用于将大语言模型适配特定任务。参数高效微调(PEFT)方法如低秩适应(LoRA)常被用来降低计算成本。PortLLM是一种无需训练、无需数据的模型适配方案,适用于持续预训练后的模型调整。尽管初始结果表明LoRA模块具有短期时间可迁移性,但其在多次持续预训练更新下的长期性能仍缺乏研究。此外,PortLLM的有效性也缺乏理论解释。本文通过(1)对Mistral、Gemma、Qwen三个基础模型在10次持续预训练步骤中进行大规模实证研究;(2)提供两种理论分析,解释为何简单方案能取得良好效果。实验发现,适配模块的可迁移性在长时间跨度下依然存在,说明基础模型定期更新时无需重复微调。理论分析表明,高维空间中向量的近正交性是时间可迁移性的关键原因。同时,该分析还揭示了损失曲面的几何特性,为不同适配方式的理论比较提供了新视角。

原文摘要 · Abstract (English)

Fine-tuning has been widely used to adapt large language models (LLMs) for domain-specific tasks. Parameter efficient fine-tuning (PEFT) methods such as low-rank adaptation (LoRA) are frequently used to reduce computational costs. PortLLM is a training-free and data-free scheme used to adapt LLMs after continual pretraining. Although the initial PortLLM results show that LoRA patches exhibit short-term temporal portability, the long-term performance of PortLLM across several updates of continual pretraining remains underexplored. Furthermore, the intriguing effectiveness of PortLLM is not well understood from a theoretical standpoint. We address these two open questions by (1) performing an extensive empirical study of the long-term temporal portability of PortLLM patches across 10 continual pretraining steps using base models Mistral, Gemma, and Qwen; and (2) offering two theoretical analyses to explain our observation that the simple PortLLM method achieves competitive performance. We find empirically that the portability persists across longer time duration, indicating that repeated fine-tuning is not required when the base model is periodically updated. We find theoretically that near-orthogonality of high-dimensional vectors is a key justification for temporal portability. Our analyses also demonstrate a geometric perspective of the loss landscape in facilitating the theoretical comparison of different adaptation options.

微调高维几何模型适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。