arXiv:2412.11087cs.IR2024-12AAAI被引 20

用大模型理解用户意图,提升图文混合检索精度

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval

  • 以大视觉语言模型为意图感知编码器,精准捕捉用户需求
  • 在三个基准上达最优效果,推理效率合理
  • 适合需要精准理解用户意图的图像检索场景

组合图像检索(CIR)旨在通过包含参考图像和描述用户意图的相对描述的多模态查询,从候选集中检索目标图像。现有方法虽尝试利用视觉语言预训练模型(VLPMs)及多种融合策略,但通常难以同时满足全面提取视觉信息和忠实遵循用户意图这两个关键要求。本文提出CIR-LVLM框架,利用大视觉语言模型(LVLM)作为强大的用户意图感知编码器,以更好满足上述需求。其核心思想是利用LVLM的高级推理与指令遵循能力,准确理解并响应用户意图。此外,设计了一种新颖的混合意图指令模块,在两个层面提供显式意图引导:(1) 任务提示明确任务要求,帮助模型在任务层面辨别用户意图;(2) 实例特定的软提示,从可学习提示池中自适应选取,相比通用提示能更准确理解单个实例的意图。CIR-LVLM在三个主流基准上达到最先进性能,且具备可接受的推理效率。本研究为CIR相关领域提供了基础性洞见。

原文摘要 · Abstract (English)

Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes the user intent. Recent studies attempt to utilize Vision-Language Pre-training Models (VLPMs) with various fusion strategies for addressing the task.However, these methods typically fail to simultaneously meet two key requirements of CIR: comprehensively extracting visual information and faithfully following the user intent. In this work, we propose CIR-LVLM, a novel framework that leverages the large vision-language model (LVLM) as the powerful user intent-aware encoder to better meet these requirements. Our motivation is to explore the advanced reasoning and instruction-following capabilities of LVLM for accurately understanding and responding the user intent. Furthermore, we design a novel hybrid intent instruction module to provide explicit intent guidance at two levels: (1) The task prompt clarifies the task requirement and assists the model in discerning user intent at the task level. (2) The instance-specific soft prompt, which is adaptively selected from the learnable prompt pool, enables the model to better comprehend the user intent at the instance level compared to a universal prompt for all instances. CIR-LVLM achieves state-of-the-art performance across three prominent benchmarks with acceptable inference efficiency. We believe this study provides fundamental insights into CIR-related fields.

图像检索大模型意图理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。