arXiv:2412.13834cs.IRcs.AI2024-12被引 3

为图文检索设计智能查询建议,让系统猜中你真正想找的图。

Maybe you are looking for CroQS: Cross-modal Query Suggestion for Text-to-Image Retrieval

  • 基于视觉一致性,建议最小微调文本以探索相关图像子集。
  • 相比初始查询,建议查询在特定性上召回率提升115%以上。
  • 适合做图文检索、交互式搜索与大模型应用的研究者参考。

查询建议是信息检索中提升系统交互性和浏览体验的重要技术。在跨模态检索中,多数研究聚焦于通过自然语言查询获取相关项,但对查询建议的研究较少。本文提出跨模态查询建议新任务,旨在基于“也许你在找……”的预设,建议最小文本修改以探索视觉一致的图像子集。为此,我们构建了专用基准 CroQS,包含初始查询、分组结果集及人工标注的建议查询。设计专门评估指标,衡量建议的代表性、聚类特异性和与原查询的相似性。将图像描述和内容摘要等领域的基线方法适配至此任务,结果显示:基于LLM和基于描述的方法虽仍远低于人类表现,但在交叉模态查询建议上表现出色,相较初始查询,聚类特异性召回率提升超115%,代表性mAP提升超52%。数据集、基线实现及实验笔记已开源。

原文摘要 · Abstract (English)

Query suggestion, a technique widely adopted in information retrieval, enhances system interactivity and the browsing experience of document collections. In cross-modal retrieval, many works have focused on retrieving relevant items from natural language queries, while few have explored query suggestion solutions. In this work, we address query suggestion in cross-modal retrieval, introducing a novel task that focuses on suggesting minimal textual modifications needed to explore visually consistent subsets of the collection, following the premise of ''Maybe you are looking for''. To facilitate the evaluation and development of methods, we present a tailored benchmark named CroQS. This dataset comprises initial queries, grouped result sets, and human-defined suggested queries for each group. We establish dedicated metrics to rigorously evaluate the performance of various methods on this task, measuring representativeness, cluster specificity, and similarity of the suggested queries to the original ones. Baseline methods from related fields, such as image captioning and content summarization, are adapted for this task to provide reference performance scores. Although relatively far from human performance, our experiments reveal that both LLM-based and captioning-based methods achieve competitive results on CroQS, improving the recall on cluster specificity by more than 115% and representativeness mAP by more than 52% with respect to the initial query. The dataset, the implementation of the baseline methods and the notebooks containing our experiments are available here: https://paciosoft.com/CroQS-benchmark/

跨模态查询建议图文检索大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。