提出SPTR方法,在不遗忘通用知识的前提下提升视觉语言模型的泛化能力。
A Similarity Paradigm Through Textual Regularization Without Forgetting
- 用最优传输做文本正则化,防止提示词过拟合。
- 设计相似性范式,提升模型在多数据集上的鲁棒性。
- 适合需要跨数据集泛化的下游任务研究者。
提示学习已成为适应预训练视觉-语言模型(VLMs)于各类下游任务的有力方法。尽管优化上下文能有效提升特定任务性能,但常导致对未见类别或分布不同数据集的泛化性能下降,原因在于文本提示易过拟合下游数据分布,从而遗忘由人工设计提示获得的通用知识。本文提出一种名为带文本正则化的相似性范式(SPTR)的新方法,实现无遗忘的提示学习。SPTR基于人工设计提示的双路径架构:1)引入最优传输作为文本正则化,精细保证人工特征与可调特征之间的逼近;2)提出自然对齐分数与对抗对齐分数的相似性范式,以持续释放多个手工提示的泛化能力。两个模块共享同一目标,旨在最大化源自多个人工提示的泛化能力。在11个数据集上,涵盖非泛化少样本学习、基类到新类泛化、跨数据集泛化及领域泛化四个代表性任务的实验表明,SPTR优于现有提示学习方法。
原文摘要 · Abstract (English)
Prompt learning has emerged as a promising method for adapting pre-trained visual-language models (VLMs) to a range of downstream tasks. While optimizing the context can be effective for improving performance on specific tasks, it can often lead to poor generalization performance on unseen classes or datasets sampled from different distributions. It may be attributed to the fact that textual prompts tend to overfit downstream data distributions, leading to the forgetting of generalized knowledge derived from hand-crafted prompts. In this paper, we propose a novel method called Similarity Paradigm with Textual Regularization (SPTR) for prompt learning without forgetting. SPTR is a two-pronged design based on hand-crafted prompts that is an inseparable framework. 1) To avoid forgetting general textual knowledge, we introduce the optimal transport as a textual regularization to finely ensure approximation with hand-crafted features and tuning textual features. 2) In order to continuously unleash the general ability of multiple hand-crafted prompts, we propose a similarity paradigm for natural alignment score and adversarial alignment score to improve model robustness for generalization. Both modules share a common objective in addressing generalization issues, aiming to maximize the generalization capability derived from multiple hand-crafted prompts. Four representative tasks (i.e., non-generalization few-shot learning, base-to-novel generalization, cross-dataset generalization, domain generalization) across 11 datasets demonstrate that SPTR outperforms existing prompt learning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。