让CLIP模型在多领域增量学习中持续进步,靠视觉与文本提示的协同作战。
ChordPrompt: Orchestrating Cross-Modal Prompt Synergy for Multi-Domain Incremental Learning in CLIP
- 设计跨模态提示,让图像和文字提示互相配合增强表达。
- 在多个数据域上实现更优的零样本泛化和下游任务性能。
- 适合需要长期适应新领域的新数据场景的开发者使用。
持续学习(CL)使预训练视觉-语言模型能够在无需全面重训练的情况下,有效适应新出现或此前未充分覆盖的数据分布,从而提升模型的适应性和效率。尽管像CLIP这样的视觉-语言模型展现出巨大潜力,但在增量学习场景下难以维持跨领域的性能。现有提示学习方法存在两大局限:1)主要针对类别增量学习,缺乏对多领域任务增量学习的具体策略;2)多数方法仅使用单模态提示,忽视了跨模态信息交互的潜在优势。为此,我们提出 extit{ChordPrompt}框架,促进视觉与文本提示之间的和谐互动。该框架引入跨模态提示以利用视觉与文本信息间的交互,并采用领域自适应文本提示,在多领域中选择合适提示实现持续适应。在多领域增量学习基准上的大量实验表明, extit{ChordPrompt}在零样本泛化能力和下游任务性能上均优于现有最先进方法。
原文摘要 · Abstract (English)
Continual learning (CL) empowers pre-trained vision-language models to adapt effectively to novel or previously underrepresented data distributions without comprehensive retraining, enhancing their adaptability and efficiency. While vision-language models like CLIP show great promise, they struggle to maintain performance across domains in incremental learning scenarios. Existing prompt learning methods face two main limitations: 1) they primarily focus on class-incremental learning scenarios, lacking specific strategies for multi-domain task incremental learning; 2) most current approaches employ single-modal prompts, neglecting the potential benefits of cross-modal information exchange. To address these challenges, we propose the \ChordPrompt framework, which facilitates a harmonious interplay between visual and textual prompts. \ChordPrompt introduces cross-modal prompts to leverage interactions between visual and textual information. Our approach also employs domain-adaptive text prompts to select appropriate prompts for continual adaptation across multiple domains. Comprehensive experiments on multi-domain incremental learning benchmarks demonstrate that \ChordPrompt outperforms state-of-the-art methods in zero-shot generalization and downstream task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。