arXiv:2605.26734cs.CV2026-05中稿 · DMLR

构建跨领域多轮图像检索数据集,解决一致性与泛化问题

CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains

论文配图:CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains
图 1 · 摘自论文原文
  • 基于CIReVL管道生成多轮查询,逐轮逼近目标图像
  • 覆盖9个领域共22,608次会话,规模超现有数据集一倍以上
  • 适合研究多轮交互式图像检索的学者与工程师

现有多轮组合图像检索(MTCIR)数据集缺乏对话历史一致性,且仅限时尚领域。为解决这些问题,我们基于FashionIQ、CIRR和CIRCO扩展构建了CIRCLED数据集。在该数据集中,每轮查询逐步接近目标图像。数据通过基于CIReVL的检索流程生成,并经多重过滤(包括检索成功率、轮数、一致性与信息冗余)确保质量。总计收集22,608个多轮会话,覆盖九个子集,规模远超多轮时尚数据集(11,505会话),在数量与泛化性上均有提升。我们还应用多种基线方法,在CIRCLED上定量评估检索精度。本工作提供了一个高质量、实用的基准,推动未来多轮图像检索研究。数据集已公开于https://huggingface.co/datasets/tk1441/CIRCLED,代码见https://github.com/mti-lab/circled。

原文摘要 · Abstract (English)

Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-historyconsistency and are restricted to the fashion domain. To address these limitations, we construct CIRCLED by extending FashionIQ, CIRR, and CIRCO. In CIRCLED, the query ateach turn progressively approaches the target image. Data are generated via a CIReVLbased retrieval pipeline and curated with multiple filters on retrieval success, turn length, consistency, and information redundancy to ensure quality. In total, we collect 22,608 multiturn sessions across nine subsets, substantially exceeding Multi-turn FashionIQ (11,505 sessions) in both scale and generality. We further apply multiple baseline methods and quantitatively assess retrieval accuracy on CIRCLED. Our work provides a practical, highquality benchmark to facilitate future research on multi-turn CIR. The dataset is publicly available at https://huggingface.co/datasets/tk1441/CIRCLED, and the code at https://github.com/mti-lab/circled.

图像检索多轮对话数据集构建跨领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。