arXiv:2604.12904cs.CV2026-04

提出新评估方法,让图像组合检索更准确可靠。

A Sanity Check on Composed Image Retrieval

  • 构建可精准控制变量的语义多样数据集FISD
  • 多轮交互式评估框架揭示模型真实表现
  • 解决旧评估中查询模糊和场景失真问题

组合图像检索(CIR)旨在根据包含参考图和相对描述的查询,找到目标图像。尽管CIR模型发展迅速,但现有基准存在查询不明确的问题(即多个候选图均满足条件),且未考虑多轮交互场景下的实际效果。为此,本文从两方面改进评估:1)提出FISD,一个全知语义多样基准,利用生成模型精确控制参考-目标图像对的变量,实现六个维度无歧义的评估;2)设计自动多轮代理评估框架,通过观察模型在多轮查询中如何调整与优化选择,更真实地评估其在实际应用中的能力。大量实验验证了新评估方法的有效性。

原文摘要 · Abstract (English)

Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image, and a relative caption that specifies the desired modification. Despite the rapid development of CIR models, their performance is not well characterized by existing benchmarks, which inherently contain indeterminate queries degrading the evaluation (i.e., multiple candidate images, rather than solely the target image, meet the query criteria), and have not considered their effectiveness in the context of the multi-round system. Motivated by this, we consider improving the evaluation procedure from two aspects: 1) we introduce FISD, a Fully-Informed Semantically-Diverse benchmark, which employs generative models to precisely control the variables of reference-target image pairs, enabling a more accurate evaluation of CIR methods across six dimensions, without query ambiguity; 2) we propose an automatic multi-round agentic evaluation framework to probe the potential of the existing models in the interactive scenarios. By observing how models adapt and refine their choices over successive rounds of queries, this framework provides a more realistic appraisal of their efficacy in practical applications. Extensive experiments and comparisons prove the value of our novel evaluation on typical CIR methods.

图像检索评估基准多轮交互生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。