构建细粒度图像检索基准,用图像编辑生成多样化查询。
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
- 通过图像编辑精准控制修改类型,合成跨类别查询
- 构建含5000个查询的EDIR基准,覆盖15个子类
- 揭示现有模型在细分类别上表现不一,适合评估多模态模型
组合图像检索(CIR)是多模态理解中的关键且复杂的任务。现有CIR基准通常查询类别有限,难以反映真实场景需求。为填补评估差距,我们利用图像编辑实现对修改类型和内容的精确控制,构建了一个可生成多样查询的流水线,并据此创建了新的细粒度CIR基准EDIR。EDIR包含5000个高质量查询,涵盖五个主类别与十五个子类别。对13个多模态嵌入模型的全面评估显示显著能力差距:即使最先进的模型(如RzenEmbed和GME)也无法在所有子类别上保持一致性能,凸显本基准的严苛性。对比分析进一步揭示现有基准的固有缺陷,如模态偏差和类别覆盖不足。此外,域内训练实验验证了本基准的可行性,通过区分可通过针对性数据解决的任务与暴露当前模型架构内在局限性的任务,明确刻画了任务挑战。
原文摘要 · Abstract (English)
Composed Image Retrieval (CIR) is a pivotal and complex task in multimodal understanding. Current CIR benchmarks typically feature limited query categories and fail to capture the diverse requirements of real-world scenarios. To bridge this evaluation gap, we leverage image editing to achieve precise control over modification types and content, enabling a pipeline for synthesizing queries across a broad spectrum of categories. Using this pipeline, we construct EDIR, a novel fine-grained CIR benchmark. EDIR encompasses 5,000 high-quality queries structured across five main categories and fifteen subcategories. Our comprehensive evaluation of 13 multimodal embedding models reveals a significant capability gap; even state-of-the-art models (e.g., RzenEmbed and GME) struggle to perform consistently across all subcategories, highlighting the rigorous nature of our benchmark. Through comparative analysis, we further uncover inherent limitations in existing benchmarks, such as modality biases and insufficient categorical coverage. Furthermore, an in-domain training experiment demonstrates the feasibility of our benchmark. This experiment clarifies the task challenges by distinguishing between categories that are solvable with targeted data and those that expose intrinsic limitations of current model architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。