低秩分解微调易遗忘,设计更新子空间可缓解灾难性遗忘
On Catastrophic Forgetting in Low-Rank Decomposition-Based Parameter-Efficient Fine-Tuning
- 通过分析更新子空间几何与参数化影响遗忘
- 张量分解法在极小参数下更好保留结构信息
- 适合需要连续学习的高效模型微调场景
基于低秩分解的参数高效微调(如LoRA)已成为适配大型预训练模型的标准方法。然而,其在序列学习中的表现——特别是灾难性遗忘问题——仍缺乏充分理解。本文通过实证研究发现,遗忘程度强烈受更新子空间的几何结构和参数化方式影响。限制更新至小而共享的矩阵子空间的方法常导致任务干扰;相比之下,张量基分解(如LoRETTA)在极小参数预算下能捕捉更丰富的结构信息,显著减轻遗忘;结构对齐的参数化方法(如WeGeFT)则有效保留预训练表示。研究强调更新子空间设计是持续学习的关键因素,并为序列场景下的高效适配策略提供实用指导。
原文摘要 · Abstract (English)
Parameter-efficient fine-tuning (PEFT) based on low-rank decomposition, such as LoRA, has become a standard for adapting large pretrained models. However, its behavior in sequential learning -- specifically regarding catastrophic forgetting -- remains insufficiently understood. In this work, we present an empirical study showing that forgetting is strongly influenced by the geometry and parameterization of the update subspace. While methods that restrict updates to small, shared matrix subspaces often suffer from task interference, tensor-based decompositions (e.g., LoRETTA) mitigate forgetting by capturing richer structural information within ultra-compact budgets, and structurally aligned parameterizations (e.g., WeGeFT) preserve pretrained representations. Our findings highlight update subspace design as a key factor in continual learning and offer practical guidance for selecting efficient adaptation strategies in sequential settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。