arXiv:2606.09648cs.DBcs.AI2026-06

构建65万条文物多模态数据集,助力跨模态错误检测与语义查询研究。

ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset

论文配图:ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset
图 1 · 摘自论文原文
  • 整合图像、文本、表格的多模态文物数据,覆盖三大博物馆真实记录。
  • 在13万条数据中发现7类领域特异性错误,如材质错位与时间错配。
  • 适合数据库、AI与文化遗产交叉研究者使用,挑战现有系统能力。

多模态数据管理已成为数据库领域的核心议题,涵盖数据集成、语义查询处理与数据质量评估。尽管关注度上升,社区仍缺乏大规模、真实世界的表格、文本与图像融合数据集。本文提出ArtiFact,一个包含651,045条来自大都会艺术博物馆、芝加哥艺术学院和荷兰国家博物馆的文物记录的多模态文化遗产数据集。我们通过两个下游任务验证其价值:在跨模态错误检测中,引入7类错误类别并标注130,209条记录,表明识别材料年代错位、时间偏差等细微领域错误仍是开放挑战;在语义查询处理中,显示当前系统难以应对文化关联性、对象类型模糊及历史术语歧义等问题。结果表明,ArtiFact可作为多模态数据管理研究的高难度基准。

原文摘要 · Abstract (English)

Multi-modal data management has emerged as a central research topic in the database community, spanning data integration, semantic query processing, and data quality assessment. Despite this growing interest, the community lacks large-scale, real-world datasets combining tables, text, and images. We present ArtiFact, a multi-modal cultural heritage dataset of 651045 museum records collected from the Metropolitan Museum of Art, the Art Institute of Chicago, and the Rijksmuseum. We demonstrate the utility of ArtiFact through two downstream tasks. For cross-modal error detection, we introduce a curated taxonomy of seven error categories injected into 130209 records and show that reliably detecting subtle domain-specific errors such as material anachronisms and temporal shifts remain an open challenge. For semantic query processing, we show that current systems struggle with queries involving cultural proximity, ambiguous object types, and historically contingent terminology. Our results position ArtiFact as a challenging benchmark for multi-modal data management research.

多模态文化遗产数据集错误检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。