提出RIFT方法,让微调时保留预训练模型的通用表征能力。
Exploring Representation Invariance in Finetuning
- 通过正交不变性设计正则化,保持微调前后表征相似性。
- 在多个低资源任务上性能不降反升,通用性显著增强。
- 适合追求迁移性能稳定的模型微调场景。
在大规模自然图像上预训练的基础模型被广泛用于各类跨域低资源下游任务,得益于其捕获的可泛化、可迁移表征。然而,研究发现这些表征在微调过程中逐渐消失,伴随模型原始泛化能力下降。本文认为,任务可有效适配而无需牺牲预训练表征的优势。为此,我们提出表示不变微调(RIFT),一种利用流形正交不变性的高效正则化方法,最大化预训练与微调后模型的表征相似性。实验表明,该方法兼容主流微调方式,在多个任务上实现相当或更优性能,并更好保留泛化能力。
原文摘要 · Abstract (English)
Foundation models pretrained on large-scale natural images are widely adapted to various cross-domain low-resource downstream tasks, benefiting from generalizable and transferable patterns captured by their representations. However, these representations are later found to gradually vanish during finetuning, accompanied by a degradation of model's original generalizability. In this paper, we argue that such tasks can be effectively adapted without sacrificing the benefits of pretrained representations. We approach this by introducing \textit{Representation Invariance FineTuning (RIFT)}, a regularization that maximizes the representation similarity between pretrained and finetuned models by leveraging orthogonal invariance of manifolds in a computationally efficient way. Experiments demonstrate that our method is compatible with mainstream finetuning methods, offering competitive or even enhanced performance and better preservation of the generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。