发现图像检索基准存在单模态捷径,高分不等于真正多模态融合。
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?

- 通过跨模型分析识别出32.2%~83.6%查询仅靠单一模态即可解决
- 人工验证后发现仅1689/4741个查询语义清晰,多数存在歧义或错配
- 剔除噪声后模型必须结合双模态信息才能准确检索,证明真融合必要性
组成式图像检索(CIR)要求模型根据参考图像和文本修改共同生成目标图像。传统认为高表现需多模态融合。本文在四个主流CIR基准和十一个通用多模态嵌入模型上发现,32.2%至83.6%的查询可仅凭单模态完成,存在普遍的单模态捷径。通过两阶段审计:首先利用跨模型分析定位可捷径解决的查询;其次对4,741个无捷径查询进行人工验证,仅1,689个为有效查询,常见问题包括修改模糊、目标错配。在该验证子集上重新评估模型,发现此时无法仅用单模态解题,成功检索必须融合双输入。尽管准确率下降,但对多模态信息的依赖显著上升。当前CIR基准混杂了可捷径、噪声及真正组合型查询,导致对模型多模态融合能力的高估。
原文摘要 · Abstract (English)
Composed Image Retrieval (CIR) is a multimodal retrieval task where a query consists of a reference image and a textual modification, and the goal is to retrieve a target image satisfying both. In principle, strong performance on CIR benchmarks is assumed to require multimodal composition, i.e., combining complementary information from reference image and textual modification. In this work, we show that this assumption does not always hold. Across four widely used CIR benchmarks and eleven Generalist Multimodal Embedding models, a large fraction of queries can be solved using a single modality (from 32.2% to 83.6%), revealing pervasive unimodal shortcuts. Thus, high CIR performance can arise from unimodal signals rather than true multimodal composition. To better understand this issue, we perform a two-stage audit. First, we identify shortcut-solvable queries through cross-model analysis. Second, we conduct human validation on 4,741 shortcut-free queries, of which only 1,689 are well-formed, with common issues including ambiguous edits and mismatched targets. Re-evaluating models on this validated subset reveals qualitatively different behaviour: queries can no longer be solved with a single modality, and successful retrieval requires combining both inputs. While accuracy decreases, reliance on multimodal information increases. Overall, current CIR benchmarks conflate shortcut-solvable, noisy, and genuinely compositional queries, leading to an overestimation of model capability in multimodal composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。