arXiv:2412.07077cs.CV2024-12被引 3

让CLIP模型在学新知识时不忘零样本能力

Retaining and Enhancing Pre-trained Knowledge in Vision-Language Models with Prompt Ensembling

  • 用分组掩码提示增强适应性,保护原模型零样本能力
  • 引入辅助提示融合新领域知识,不破坏原有表征
  • 集成学习策略提升跨数据集迁移表现,适合领域适配场景

视觉语言模型(如对比语言-图像预训练模型CLIP)通过强大的零样本学习能力推动了机器学习发展,使模型能理解未见过的数据而无需特定任务训练。然而,在保持零样本能力的同时将特定领域知识融入CLIP仍面临挑战。为此,本文提出一种新型提示集成方法——分组提示集成(Group-wise Prompt Ensemble, GPE)。该方法通过三种策略实现:使用掩码注意力的提示分组以优化适应性并保护零样本能力;引入辅助提示实现新领域知识的无缝整合而不干扰原始模型表征;采用集成学习策略有效融合原始与新知识。通过严格的实验,包括更具挑战性的跨数据集迁移评估,GPE方法重新定义了视觉语言模型在适应性与效率方面的基准表现,在多种场景下超越现有模型。

原文摘要 · Abstract (English)

The advancement of vision-language models, particularly the Contrastive Language-Image Pre-training (CLIP) model, has revolutionized the field of machine learning by enabling robust zero-shot learning capabilities. These capabilities allow models to understand and respond to previously unseen data without task-specific training. However, adapting CLIP to integrate specialized knowledge from various domains while retaining its zero-shot capabilities remains a significant challenge. To address this, we introduce a novel prompt ensemble learning approach called Group-wise Prompt Ensemble (GPE). This method aims to enhance CLIP's zero-shot capabilities by incorporating new domain knowledge while improving its adaptability and robustness against data distribution shifts. Our approach hinges on three main strategies: prompt grouping with masked attention to optimize CLIP's adaptability while safeguarding its zero-shot capabilities; the incorporation of auxiliary prompts for the seamless integration of new domain insights without disrupting the original model's representation; and an ensemble learning strategy that effectively merges original and new knowledge. Through rigorous experimentation, including more challenging cross-dataset transfer evaluations, our GPE method redefines the benchmarks for the adaptability and efficiency of vision-language models, surpassing existing models across various scenarios.

视觉语言模型提示工程知识保留零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。