用大模型增强提示学习,让CLIP理解图像中的动作细节。
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching
- 引入大语言模型生成动作相关知识,设计双提示捕捉语义组合与状态因果。
- 在COCO和Flickr30K上提升图像文本匹配准确率,显著优于基线方法。
- 适合关注视觉-语言细粒度对齐与动作理解的研究者。
受大规模对比视觉-语言预训练模型(如CLIP)驱动,图像-文本匹配任务在表征学习方面取得了显著进展。然而,由于仅依赖图像级视觉-语言对齐,CLIP难以理解物体属性及物体间空间关系等细粒度信息。近期工作尝试通过提示学习引入结构化视觉表征以实现物体级对齐,但仍未具备感知动作的能力——而动作对描述对象状态或关系至关重要。为此,我们提出一种基于大语言模型(LLM)增强的动作感知多模态提示调优方法,将由大语言模型生成的动作相关外部知识融入模型。具体地,设计动作三元组提示与动作状态提示,挖掘大语言模型中隐含的组合语义知识与状态因果知识;随后提出自适应交互模块,基于动作感知提示知识聚合注意力视觉特征,构建判别性且动作感知的视觉表征,进一步提升性能。在两个基准数据集(COCO与Flickr30K)上的全面实验验证了该方法的有效性。
原文摘要 · Abstract (English)
Driven by large-scale contrastive vision-language pre-trained models such as CLIP, recent advancements in the image-text matching task have achieved remarkable success in representation learning. Due to image-level visual-language alignment, CLIP falls short in understanding fine-grained details such as object attributes and spatial relationships between objects. Recent efforts have attempted to compel CLIP to acquire structured visual representations by introducing prompt learning to achieve object-level alignment. While achieving promising results, they still lack the capability to perceive actions, which are crucial for describing the states or relationships between objects. Therefore, we propose to endow CLIP with fine-grained action-level understanding by introducing an LLM-enhanced action-aware multi-modal prompt-tuning method, incorporating the action-related external knowledge generated by large language models (LLMs). Specifically, we design an action triplet prompt and an action state prompt to exploit compositional semantic knowledge and state-related causal knowledge implicitly stored in LLMs. Subsequently, we propose an adaptive interaction module to aggregate attentive visual features conditioned on action-aware prompted knowledge for establishing discriminative and action-aware visual representations, which further improves the performance. Comprehensive experimental results on two benchmark datasets demonstrate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。