构建首个面向文博机构的图像组合检索数据集,支持图文混合查询。
EUFCC-CIR: a Composed Image Retrieval Dataset for GLAM Collections
- 基于34万标注图像构建18万组图文组合检索三元组
- 支持以图像+文本描述属性变换的复杂查询任务
- 适用于数字人文与跨模态检索研究者
人工智能与数字人文的交叉推动了文化遗产研究的深度与广度。本文提出EUFCC-CIR数据集,专为美术馆、图书馆、档案馆和博物馆(GLAM)收藏中的组合图像检索(CIR)设计。该数据集建立在EUFCC-340K图像标注数据集之上,包含超过18万条标注的CIR三元组。每个三元组由多模态查询(一张输入图像+一段描述目标属性变换的简短文本)和一组相关目标图像构成。EUFCC-CIR填补了数字人文领域专用CIR资源的空白。我们通过对比其他现有CIR数据集并评估多个零样本CIR基线模型性能,验证了该数据集的价值。
原文摘要 · Abstract (English)
The intersection of Artificial Intelligence and Digital Humanities enables researchers to explore cultural heritage collections with greater depth and scale. In this paper, we present EUFCC-CIR, a dataset designed for Composed Image Retrieval (CIR) within Galleries, Libraries, Archives, and Museums (GLAM) collections. Our dataset is built on top of the EUFCC-340K image labeling dataset and contains over 180K annotated CIR triplets. Each triplet is composed of a multi-modal query (an input image plus a short text describing the desired attribute manipulations) and a set of relevant target images. The EUFCC-CIR dataset fills an existing gap in CIR-specific resources for Digital Humanities. We demonstrate the value of the EUFCC-CIR dataset by highlighting its unique qualities in comparison to other existing CIR datasets and evaluating the performance of several zero-shot CIR baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。