用生成图像反馈帮助用户逐步完善心中所想的图片搜索。
GenIR: Generative Visual Feedback for Mental Image Retrieval

- 用扩散模型生成视觉反馈,让系统理解更直观可感知。
- 在多轮交互中显著提升检索准确率,优于现有方法。
- 适合需要反复调整搜索意图的智能图像检索场景。
视觉语言模型在文本到图像检索任务上表现优异,但在真实场景中的应用仍面临挑战。现实中,人类搜索往往不是一次完成,而是通过多轮交互,基于脑海中的模糊或清晰图像记忆逐步修正目标。为此,本文提出心理图像检索(MIR)任务,聚焦于用户通过与搜索系统多轮互动来逼近心中所想图像的现实场景。现有方法依赖抽象或间接的语言反馈,易产生歧义或无效引导。为此,本文提出GenIR,一种基于扩散模型生成视觉反馈的多轮检索范式,将系统理解转化为可直观感知的合成图像,使用户能有效调整查询。我们还构建了全自动生成的高质量多轮MIR数据集。实验表明,GenIR在MIR场景下显著优于现有交互式方法。该工作确立了新任务、提供了数据集与有效方法,为后续研究奠定基础。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have shown strong performance on text-to-image retrieval benchmarks. However, bridging this success to real-world applications remains a challenge. In practice, human search behavior is rarely a one-shot action. Instead, it is often a multi-round process guided by clues in mind. That is, a mental image ranging from vague recollections to vivid mental representations of the target image. Motivated by this gap, we study the task of Mental Image Retrieval (MIR), which targets the realistic yet underexplored setting where users refine their search for a mentally envisioned image through multi-round interactions with an image search engine. Central to successful interactive retrieval is the capability of machines to provide users with clear, actionable feedback; however, existing methods rely on indirect or abstract verbal feedback, which can be ambiguous, misleading, or ineffective for users to refine the query. To overcome this, we propose GenIR, a generative multi-round retrieval paradigm leveraging diffusion-based image generation to explicitly reify the AI system's understanding at each round. These synthetic visual representations provide clear, interpretable feedback, enabling users to refine their queries intuitively and effectively. We further introduce a fully automated pipeline to generate a high-quality multi-round MIR dataset. Experimental results demonstrate that GenIR significantly outperforms existing interactive methods in the MIR scenario. This work establishes a new task with a dataset and an effective generative retrieval method, providing a foundation for future research in this direction
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。