用视觉选项替代文字提问,更准更快地解决图像搜索歧义
Show, Don't Ask: Generative Visual Disambiguation for Composed Image Retrieval with Turn-Valid Coverage

- 用户从一组真实候选图中选最接近目标的,直接反馈意图
- 多轮交互下仍保持95%覆盖率,优于仅首轮有效的旧方法
- 适合需要分辨视角、细节属性的复杂图像搜索场景
组合图像检索(CIR)通过参考图和文本修改查找目标图,但查询常存在多种可能,导致意图模糊。现有方法依赖可信预测估计模糊性并提出文字问题澄清,但存在两个缺陷:覆盖率仅在第一轮有效,且文字问题难以解决外观、属性或视角等细粒度差异。本文提出CLARA框架,通过展示一组视觉候选图来澄清歧义,用户只需选择最接近目标的原型图,从而提供直接视觉反馈,避免依赖模型对回答的预测。为在多轮交互中维持有效的覆盖率,CLARA基于用户选择的似然比重新校准,同时约束候选图来自真实数据集并贴合实际图像,防止生成图人为提升覆盖率。在开放域和时尚基准测试中,CLARA达到单轮先进水平的检索性能,多轮下保持名义覆盖率,并以更少轮次找到目标。当歧义涉及视角或细粒度属性时,其优势尤为显著。
原文摘要 · Abstract (English)
Composed image retrieval (CIR) uses a reference image and a text modification to search for a target image. However, such queries often describe several possible images rather than one exact target, making the user's intent ambiguous. Recent methods address this by using conformal prediction to estimate ambiguity and by asking users clarifying text questions. However, these methods have two limitations: their coverage guarantee only holds at the first interaction, and text questions are often insufficient for resolving fine-grained visual differences such as appearance, attributes, or viewpoint. We propose CLARA, a clarification framework that resolves ambiguity by showing users a small panel of visual alternatives. Instead of answering text questions, the user simply selects the prototype image closest to the intended target. This provides a direct visual signal and avoids relying on a model to predict the user's answer. To maintain valid conformal guarantees across multiple interaction rounds, CLARA reweights calibration using the likelihood ratio induced by the user's selection. The displayed prototypes are also constrained to represent the current candidate set and are snapped to real corpus images, ensuring that generated images cannot artificially improve coverage. Experiments on open-domain and fashion benchmarks show that CLARA matches single-turn state-of-the-art retrieval performance, maintains nominal coverage across interaction rounds, and finds the intended target in fewer rounds than strong text-question baselines. Its advantage is especially clear when ambiguity involves viewpoint or fine-grained attributes, where visual clarification is more effective than textual questioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。