构建首个藏传佛画的草图+文本组合检索数据集,支持细粒度文化语义理解。
A Sketch+Text Composed Image Retrieval Dataset for Thangka
- 基于草图与多层级文本描述构建组合查询,融合结构意图与语义信息
- 包含2287幅高质量唐卡图像,每张配手绘草图与三级语义文本
- 适用于文化遗产、知识密集型视觉域的多模态检索研究
组合图像检索(CIR)通过结合多种查询模态实现图像搜索,但现有基准主要集中于通用图像领域,依赖短文本修改的参考图像,难以支持需要细粒度语义推理、结构化视觉理解及领域专有知识的检索场景。本文提出CIRThan,一个面向唐卡图像的草图+文本组合图像检索数据集,该领域具有复杂结构、密集符号元素和依赖领域语义规范的文化特征。CIRThan包含2,287幅高质量唐卡图像,每张图像均配有手工绘制的草图和三级语义层次的文本描述,支持联合表达结构意图与多层次语义指定的组合查询。我们提供了标准化数据划分、全面的数据集分析以及代表性监督与零样本CIR方法的基准评估。实验结果表明,现有针对通用图像设计的CIR方法在对齐草图抽象与层级文本语义与细粒度唐卡图像方面表现不佳,尤其在缺乏领域内监督时更为明显。我们认为CIRThan为推进草图+文本CIR、层级语义建模及文化遗产等知识密集型视觉领域的多模态检索提供了重要基准。数据集已公开于https://github.com/jinyuxu-whut/CIRThan。
原文摘要 · Abstract (English)
Composed Image Retrieval (CIR) enables image retrieval by combining multiple query modalities, but existing benchmarks predominantly focus on general-domain imagery and rely on reference images with short textual modifications. As a result, they provide limited support for retrieval scenarios that require fine-grained semantic reasoning, structured visual understanding, and domain-specific knowledge. In this work, we introduce CIRThan, a sketch+text Composed Image Retrieval dataset for Thangka imagery, a culturally grounded and knowledge-specific visual domain characterized by complex structures, dense symbolic elements, and domain-dependent semantic conventions. CIRThan contains 2,287 high-quality Thangka images, each paired with a human-drawn sketch and hierarchical textual descriptions at three semantic levels, enabling composed queries that jointly express structural intent and multi-level semantic specification. We provide standardized data splits, comprehensive dataset analysis, and benchmark evaluations of representative supervised and zero-shot CIR methods. Experimental results reveal that existing CIR approaches, largely developed for general-domain imagery, struggle to effectively align sketch-based abstractions and hierarchical textual semantics with fine-grained Thangka images, particularly without in-domain supervision. We believe CIRThan offers a valuable benchmark for advancing sketch+text CIR, hierarchical semantic modeling, and multimodal retrieval in cultural heritage and other knowledge-specific visual domains. The dataset is publicly available at https://github.com/jinyuxu-whut/CIRThan.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。