arXiv:2505.23543cs.CV2025-05

用视觉模型+知识图谱自动丰富文物元数据,提升文化遗产数字化利用效率。

Position Paper: Metadata Enrichment Model: Integrating Neural Networks and Semantic Knowledge Graphs for Cultural Heritage Applications

  • 融合视觉模型与知识图谱,动态识别印章内文字、图像等嵌套特征
  • 在105页古籍上验证,可提升元数据的结构化与语义关联度
  • 适合博物馆、档案馆等机构开展智能文物管理与跨库协作

文化遗产数字化带来了新研究方向,但元数据匮乏严重制约其可访问性、互操作性与跨机构协作。近年来YOLOv11、Detectron2等神经网络革新了视觉分析,但在手稿、早期印刷品等特定领域仍受限于缺乏结构特征提取与语义互通的方法。本文提出元数据增强模型(MEM),融合微调计算机视觉模型、大语言模型(LLMs)与结构化知识图谱,实现对数字藏品元数据的深度增强。其核心创新为多层视觉机制(MVM),通过迭代方式动态检测嵌套特征,如印章中的文字或邮戳中的图像。我们以雅盖隆数字图书馆的105页古籍数据集为例,发布人工标注数据集并验证该方法。讨论了实际应用中需进行领域微调、遵循链接数据标准及计算成本等问题。MEM具备良好扩展性,为人工智能与语义网技术在文化遗产领域的实践提供可行路径。

原文摘要 · Abstract (English)

The digitization of cultural heritage collections has opened new directions for research, yet the lack of enriched metadata poses a substantial challenge to accessibility, interoperability, and cross-institutional collaboration. In several past years neural networks models such as YOLOv11 and Detectron2 have revolutionized visual data analysis, but their application to domain-specific cultural artifacts - such as manuscripts and incunabula - remains limited by the absence of methodologies that address structural feature extraction and semantic interoperability. In this position paper, we argue, that the integration of neural networks with semantic technologies represents a paradigm shift in cultural heritage digitization processes. We present the Metadata Enrichment Model (MEM), a conceptual framework designed to enrich metadata for digitized collections by combining fine-tuned computer vision models, large language models (LLMs) and structured knowledge graphs. The Multilayer Vision Mechanism (MVM) appears as the key innovation of MEM. This iterative process improves visual analysis by dynamically detecting nested features, such as text within seals or images within stamps. To expose MEM's potential, we apply it to a dataset of digitized incunabula from the Jagiellonian Digital Library and release a manually annotated dataset of 105 manuscript pages. We examine the practical challenges of MEM's usage in real-world GLAM institutions, including the need for domain-specific fine-tuning, the adjustment of enriched metadata with Linked Data standards and computational costs. We present MEM as a flexible and extensible methodology. This paper contributes to the discussion on how artificial intelligence and semantic web technologies can advance cultural heritage research, and also use these technologies in practice.

文化遗产元数据增强知识图谱视觉分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。