通过知识蒸馏提升视觉语言模型在未见类别上的泛化能力
MoPD: Mixture-of-Prompts Distillation for Vision-Language Models
- 用人工提示作为教师,蒸馏知识到可学习的软提示中
- 在未见类别上性能显著优于现有方法
- 适合需要强泛化能力的下游任务场景
软提示学习在适配视觉语言模型(VLMs)至下游任务时表现良好。然而,现有方法存在对已见类别过拟合、在未见类别上性能下降的问题,根源在于训练数据对已见类别的固有偏倚。为此,我们提出一种新型软提示学习方法——混合提示蒸馏(MoPD),通过将人工设计的硬提示(教师提示)中的有用知识迁移到可学习的软提示(学生提示)中,有效提升软提示在未见类别上的泛化能力。此外,该方法引入门控网络,自动选择用于蒸馏的硬提示。大量实验表明,所提方法在未见类别上显著优于当前最优基线。
原文摘要 · Abstract (English)
Soft prompt learning methods are effective for adapting vision-language models (VLMs) to downstream tasks. Nevertheless, empirical evidence reveals a tendency of existing methods that they overfit seen classes and exhibit degraded performance on unseen classes. This limitation is due to the inherent bias in the training data towards the seen classes. To address this issue, we propose a novel soft prompt learning method, named Mixture-of-Prompts Distillation (MoPD), which can effectively transfer useful knowledge from hard prompts manually hand-crafted (a.k.a. teacher prompts) to the learnable soft prompt (a.k.a. student prompt), thereby enhancing the generalization ability of soft prompts on unseen classes. Moreover, the proposed MoPD method utilizes a gating network that learns to select hard prompts used for prompt distillation. Extensive experiments demonstrate that the proposed MoPD method outperforms state-of-the-art baselines especially on on unseen classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。