arXiv:2502.20292cs.CVcs.LG2025-02ICCV被引 2

让视觉提示随图像动态调整,提升零样本组合识别能力

Visual Adaptive Prompting for Compositional Zero-Shot Learning

  • 根据图像特征动态选择最相关的属性和物体提示
  • 在三个基准上达到当前最优性能,闭/开世界场景均有效
  • 适合需要灵活组合推理的视觉语言模型研究者

视觉语言模型(VLMs)在联合学习视觉与文本表示方面表现出色,是实现组合零样本学习(CZSL)的强大工具。CZSL要求模型能泛化到训练中未显式出现过的视觉要素(如属性与对象)的新组合。现有提示方法多聚焦于文本编码器输入的修改,使用固定提示,难以适应不同视觉上下文。为此,本文提出视觉自适应提示系统(VAPS),通过可学习的视觉提示库与基于相似性的检索机制,在VLM框架内弥合语义与视觉特征的差距。该方法引入动态视觉提示库,依据图像特征选择最相关属性和物体提示,并设计视觉提示适配器以学习更具泛化性的嵌入空间。在三个CZSL基准上,涵盖闭世界与开世界场景的实验表明,本方法取得当前最优结果。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated impressive multimodal capabilities in learning joint representations of visual and textual data, making them powerful tools for tasks such as Compositional Zero-Shot Learning (CZSL). CZSL requires models to generalize to novel combinations of visual primitives--such as attributes and objects--that were not explicitly encountered during training. Recent works in prompting for CZSL have focused on modifying inputs for the text encoder, often using static prompts that do not change across varying visual contexts. However, these approaches struggle to fully capture varying visual contexts, as they focus on text adaptation rather than leveraging visual features for compositional reasoning. To address this, we propose a Visual Adaptive Prompting System (VAPS) that leverages a learnable visual prompt repository and similarity-based retrieval mechanism within the framework of VLMs to bridge the gap between semantic and visual features. Our method introduces a dynamic visual prompt repository mechanism that selects the most relevant attribute and object prompts based on the visual features of the image. Our proposed system includes a visual prompt adapter that encourages the model to learn a more generalizable embedding space. Experiments on three CZSL benchmarks, across both closed and open-world scenarios, demonstrate state-of-the-art results.

视觉语言模型零样本学习提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。