构建中文文化遗产多模态数据集并提出新方法提升图文检索精度
Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution
- 基于古丝绸与敦煌壁画文档构建5726对图文数据集
- 新方法LACLIP在细粒度语义匹配上显著优于现有模型
- 适合研究文化遗产数字化与跨模态检索的学者使用
中国拥有悠久丰富的历史,涵盖大量多模态文化资源,如丝织图案、敦煌壁画及其相关历史叙述。跨模态检索在理解与解读中国文化遗产中至关重要,能连接视觉与文本模态,实现精准的文图互检。然而,尽管多模态研究日益兴起,针对中国文化遗产的专用数据集仍十分匮乏,制约了该领域跨模态学习模型的发展与评估。为此,我们提出了一个名为CulTi的多模态数据集,包含从两套专业文献中提取的5,726对图像-文本数据,分别涉及古代中国丝绸与敦煌壁画。相较于现有通用多模态数据集,CulTi在复杂装饰纹样与专业文本描述之间的局部对齐方面更具挑战性。为应对这一难题,我们提出LACLIP——一种无需训练的局部对齐策略,基于微调后的中文CLIP,在推理阶段通过加权相似度计算增强全局文本描述与局部视觉区域的对齐。在CulTi上的实验结果表明,LACLIP在跨模态检索任务中显著优于现有模型,尤其在处理中国文化遗产中的细粒度语义关联方面表现突出。
原文摘要 · Abstract (English)
China has a long and rich history, encompassing a vast cultural heritage that includes diverse multimodal information, such as silk patterns, Dunhuang murals, and their associated historical narratives. Cross-modal retrieval plays a pivotal role in understanding and interpreting Chinese cultural heritage by bridging visual and textual modalities to enable accurate text-to-image and image-to-text retrieval. However, despite the growing interest in multimodal research, there is a lack of specialized datasets dedicated to Chinese cultural heritage, limiting the development and evaluation of cross-modal learning models in this domain. To address this gap, we propose a multimodal dataset named CulTi, which contains 5,726 image-text pairs extracted from two series of professional documents, respectively related to ancient Chinese silk and Dunhuang murals. Compared to existing general-domain multimodal datasets, CulTi presents a challenge for cross-modal retrieval: the difficulty of local alignment between intricate decorative motifs and specialized textual descriptions. To address this challenge, we propose LACLIP, a training-free local alignment strategy built upon a fine-tuned Chinese-CLIP. LACLIP enhances the alignment of global textual descriptions with local visual regions by computing weighted similarity scores during inference. Experimental results on CulTi demonstrate that LACLIP significantly outperforms existing models in cross-modal retrieval, particularly in handling fine-grained semantic associations within Chinese cultural heritage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。