用对话重写提升多模态图像检索的准确率
ReCQR: Incorporating conversational query rewriting to improve Multimodal Image Retrieval
- 基于对话历史重写用户查询,生成更精准的检索语句
- 在7000条高质量对话数据上验证,显著提升检索精度
- 适合研究多模态交互与自然语言理解的学者
随着多模态学习的发展,图像检索在连接视觉信息与自然语言查询方面发挥着关键作用。现有图像检索模型难以处理长文本和模糊表达。为此,本文将对话式查询重写(CQR)引入图像检索领域,构建了专门的多轮对话重写数据集。基于完整的对话历史,CQR将用户的最终查询重写为简洁且语义完整的形式,更利于检索。具体而言,首先利用大语言模型(LLMs)大规模生成重写候选,再通过LLM作为裁判结合人工审核,筛选出约7000条高质量多模态对话,形成ReCQR数据集。随后在该数据集上对多个SOTA多模态模型进行基准测试,评估其在图像检索中的表现。实验结果表明,CQR不仅显著提升了传统图像检索模型的准确性,也为多模态系统中用户查询建模提供了新思路与洞见。
原文摘要 · Abstract (English)
With the rise of multimodal learning, image retrieval plays a crucial role in connecting visual information with natural language queries. Existing image retrievers struggle with processing long texts and handling unclear user expressions. To address these issues, we introduce the conversational query rewriting (CQR) task into the image retrieval domain and construct a dedicated multi-turn dialogue query rewriting dataset. Built on full dialogue histories, CQR rewrites users' final queries into concise, semantically complete ones that are better suited for retrieval. Specifically, We first leverage Large Language Models (LLMs) to generate rewritten candidates at scale and employ an LLM-as-Judge mechanism combined with manual review to curate approximately 7,000 high-quality multimodal dialogues, forming the ReCQR dataset. Then We benchmark several SOTA multimodal models on the ReCQR dataset to assess their performance on image retrieval. Experimental results demonstrate that CQR not only significantly enhances the accuracy of traditional image retrieval models, but also provides new directions and insights for modeling user queries in multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。