通过对话迭代优化,让图文检索更准更智能。
DIR-TIR: Dialog-Iterative Refinement for Text-to-Image Retrieval
- 用对话逐步追问用户,精准提取目标图像描述。
- 结合图像反馈,主动修正视觉与语义差距,提升检索准确率。
- 适合需要精细控制的图像搜索场景,如设计、创作辅助。
本文针对交互式对话式文本到图像检索任务,提出DIR-TIR框架,通过对话精炼模块和图像精炼模块协同工作,逐步优化目标图像搜索。对话精炼模块主动向用户提问,获取关键信息并生成更精确的图像描述;图像精炼模块则识别生成图像与用户意图之间的感知差异,有针对性地缩小视觉-语义差距。基于多轮对话,该方法相比传统单次查询方式具备更强可控性与容错能力,显著提升目标图像命中率。在多个图像数据集上的实验证明,该对话驱动方法显著优于仅依赖初始描述的基线模型,且模块协同实现了更高的检索精度与更优的交互体验。
原文摘要 · Abstract (English)
This paper addresses the task of interactive, conversational text-to-image retrieval. Our DIR-TIR framework progressively refines the target image search through two specialized modules: the Dialog Refiner Module and the Image Refiner Module. The Dialog Refiner actively queries users to extract essential information and generate increasingly precise descriptions of the target image. Complementarily, the Image Refiner identifies perceptual gaps between generated images and user intentions, strategically reducing the visual-semantic discrepancy. By leveraging multi-turn dialogues, DIR-TIR provides superior controllability and fault tolerance compared to conventional single-query methods, significantly improving target image hit accuracy. Comprehensive experiments across diverse image datasets demonstrate our dialogue-based approach substantially outperforms initial-description-only baselines, while the synergistic module integration achieves both higher retrieval precision and enhanced interactive experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。