arXiv:2410.03189cs.CV2024-10被引 4

让视觉语言模型同时具备强任务性能和跨类别泛化能力

Generalizable Prompt Tuning for Vision-Language Models

  • 将可学习提示与手工提示视为文本模态的双视角,最大化互信息融合
  • 引入类别级视觉增强生成更丰富提示,显著提升对未见类别的鲁棒性
  • 适合需要兼顾精准任务表现与泛化能力的研究者使用

针对视觉语言模型如CLIP的提示调优问题,现有方法在任务特定性能与泛化能力间存在权衡。手工模板提示虽泛化性强但下游表现差,可学习软提示性能好却缺乏泛化性。本文提出将软提示与手工提示视为文本模态的双视角,通过最大化其互信息实现任务特异性与通用语义信息的融合;同时引入类别级视觉模态增强,生成更具表达力的提示。在多个基准测试中,该方法在任务特定性能与泛化能力上均取得竞争力结果。

原文摘要 · Abstract (English)

Prompt tuning for vision-language models such as CLIP involves optimizing the text prompts used to generate image-text pairs for specific downstream tasks. While hand-crafted or template-based prompts are generally applicable to a wider range of unseen classes, they tend to perform poorly in downstream tasks (i.e., seen classes). Learnable soft prompts, on the other hand, often perform well in downstream tasks but lack generalizability. Additionally, prior research has predominantly concentrated on the textual modality, with very few studies attempting to explore the prompt's generalization potential from the visual modality. Keeping these limitations in mind, we investigate how to prompt tuning to obtain both a competitive downstream performance and generalization. The study shows that by treating soft and hand-crafted prompts as dual views of the textual modality, and maximizing their mutual information, we can better ensemble task-specific and general semantic information. Moreover, to generate more expressive prompts, the study introduces a class-wise augmentation from the visual modality, resulting in significant robustness to a wider range of unseen classes. Extensive evaluations on several benchmarks report that the proposed approach achieves competitive results in terms of both task-specific performance and general abilities.

提示调优多模态泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。